Risk wording dilution is the weakening or obfuscation of material risk language in public disclosures, whether in 10-K risk factors, earnings-call transcripts, or press releases, that reduces investor clarity and raises mispricing and governance risk. The SEC's own guidance on plain-English disclosure and a 2024 structural-detection framework for 10-K narrative manipulation both point to the same conclusion: dilution is measurable, not merely a matter of tone. Lacuna Index treats it as a monitoring trigger, not a stylistic quirk.
TL;DR:
- Hedge proliferation and sentence inflation are key visual signals of risk wording dilution that can be tracked through hedge term counts and sentence length analysis.
- Semantic embedding methods reveal structural risks and language shifts over multiple years, capturing restructuring that word counts and readability scores may miss.
- Combining multiple metrics such as semantic distance, hedge density, and textual similarity produces more reliable detection of dilution trends than relying on a single measure.
- Flags should be prioritized based on materiality, corroboration, and novelty, with evidence logs providing traceability for governance or journalistic review.
- Regular monitoring using multi-year baselines and sector benchmarks helps detect cumulative language dilution, which can then be addressed through engagement rather than enforcement.
Table of Contents
- Spotting Risk Wording Dilution in the Wild
- Word Counts Versus Structural Semantics: Which Method Actually Works
- Deciding Which Flags Deserve a Closer Look
- What Lacuna Index's Evidence Pipeline Actually Surfaces
- Building a Monitoring Pipeline Analysts Can Actually Run
- Before-and-After: What Dilution Looks Like on the Page
- Why Dilution Happens: The Usual Culprits
- Measuring How Much Dilution Has Accumulated Over Time
- When Diluted Disclosures Become a Compliance Problem
- Bringing Detection Into an Existing Review Workflow
- Detection in Practice: What a Caught Case Looks Like
- Turn Detection Into a Standing Practice, Not a One-Off Audit
- Why Explainability Beats Alerts
- Sources
Spotting Risk Wording Dilution in the Wild
Dilution rarely announces itself. It shows up as a slow drift in the texture of a filing, one that a careless reader skims past and a forensic reader catches on the second pass. Five observable patterns cover most of what analysts encounter in practice.
Hedge proliferation is the accumulation of qualifying language, "may," "could," "in certain circumstances," "under some conditions", stacked densely enough that a sentence no longer commits to a claim. Counting hedge terms per hundred words, then tracking that ratio release over release, turns a vague impression into a data point.
Sentence inflation works alongside hedging. The Harvard Law School Forum's analysis of SEC risk-factor rules points to an average sentence length within a recommended Plain English benchmark; filings that exceed recommended sentence lengths tend to bury the actual exposure inside subordinate clauses.
Beyond sentence mechanics, three structural signals matter just as much:
- Generic headings replacing named, specific risk categories ("Various Factors May Affect Our Business" instead of a named commodity, customer, or regulatory exposure)
- Loss of named specifics, where a supplier's name, a dollar exposure, or a jurisdiction disappears from one filing cycle to the next without explanation
- Boilerplate drift, meaning high year-over-year textual similarity in sections that should reflect a changing business, often accelerated by generative-AI drafting tools
The highest-priority flag combines all four with a fifth: a tone-fundamentals gap, where confident, upbeat narrative language sits next to deteriorating margins, rising churn, or a fresh restatement. The 2024 narrative-manipulation framework identifies this gap as one of the more reliable signals of engineered disclosure, and it's the pattern Lacuna Index weights most heavily in its own scoring.
Word Counts Versus Structural Semantics: Which Method Actually Works
Dictionary-based methods answer a narrower question than most analysts assume. The Loughran-McDonald financial sentiment word lists count how many negative, uncertain, or litigious terms appear in a filing, but they cannot tell you whether a sentence was restructured to soften an admission without changing its vocabulary at all. A company can remove every flagged word from a risk factor while making the actual disclosure less specific, and a pure word-count check will show improvement.
Readability metrics move a step closer to what the SEC's Plain English initiative was built for: shorter sentences, purposeful headings drawn from an established risk taxonomy, and concise summaries when a risk-factor section runs long. These metrics correlate with something real. Research on European-listed banks found that less-readable risk disclosures tend to accompany more aggressive discretionary accruals, consistent with a management-obfuscation pattern rather than mere writing style.
Sentence embeddings solve the blind spot that word counts and readability scores share. By representing each sentence as a vector in semantic space, an analyst can measure how far a company's risk narrative has moved from its own historical baseline, or from its peer group, even when the vocabulary looks unchanged. A 2026 preprint on structural deviation in 10-K risk factors demonstrates that this semantic-dispersion approach catches restructuring that dictionary methods miss entirely.
Here is the recommended sequence for building a detection pipeline:
- Run readability and hedge-density checks first, since they're cheap and catch the obvious cases.
- Generate sentence-level embeddings for each risk-factor section and compare against a rolling historical baseline for the same issuer and its peer group.
- Feed the semantic-distance scores into an anomaly detector, an Isolation Forest model tends to perform well here, tuned for stability rather than sensitivity.
- Decompose flagged anomalies with a feature-level explainer such as SHAP so the output reads as evidence, not a black-box score.
- Cross-check any flag against a baseline stability test before escalating, since a single noisy quarter shouldn't trigger the same response as a persistent trend.
Pro Tip: Run your anomaly detector against at least three prior fiscal years before trusting its first flag. A one-year baseline mistakes normal drafting turnover for genuine dilution far more often than analysts expect.
No single metric should carry a triage decision alone. Combining hedge density, sentence length, semantic dispersion, and structural similarity produces a more robust detection signal than any single feature, and it gives you multiple independent reasons to trust or discard a flag.
Deciding Which Flags Deserve a Closer Look
A monitoring pipeline that flags everything is functionally identical to one that flags nothing. Triage has to run on materiality, timing, and corroboration, in that order.
Materiality filters come first. A hedge-density spike in a risk factor tied to a small, non-core business line matters less than the same spike in the section covering the company's primary revenue driver. Weight flags by exposure size, by recent governance signals (departing CFOs, delayed filings, auditor changes), and by any recent deterioration in fundamentals.
Novelty versus boilerplate separates noise from signal. A sudden new risk factor appearing for the first time is expected and often benign. A sudden new generic risk factor appearing alongside unusually high similarity to the prior year's language elsewhere in the filing is a different story. That combination, novelty in one place and static boilerplate everywhere else, tends to indicate a company reaching for cover rather than disclosing a genuine new exposure.
Corroborating checks turn a flag into a case worth escalating:
- Discretionary accruals or accounting quality metrics moving in the same direction as the readability decline
- Recent restatements or late-filing notices
- Enforcement history for the company or its auditor
- Unusual market reaction in the days surrounding the filing date
A simple rubric works well in practice: escalate flags with high materiality and strong corroboration, report for review flags with either but not both, and watch flags that show textual anomaly with no corroborating signal yet. Analysts can build a fuller taxonomy of what counts as a material disclosure omission to standardize this triage across a coverage universe.
What Lacuna Index's Evidence Pipeline Actually Surfaces
Lacuna Index builds its narrative-versus-delivery scoring on a pipeline that starts with the same public filings any analyst can access: 10-Ks, proxy statements, earnings-call transcripts, and press releases. The text is parsed into discrete sections, risk factors separated from MD&A, forward-looking statements separated from historical results, then scored against the company's own execution record and against sector peers. The output classifies companies into archetypes such as earned, borrowed, or undervalued, depending on how closely communicated narrative tracks actual delivery.
A representative pattern looks like this: a company's earnings-call language stays confident and forward-looking across three consecutive quarters, while its 10-K risk-factor section simultaneously adds three new generic risk categories and shows a measurable jump in semantic distance from its own two-year baseline. Decomposed feature-by-feature, the anomaly traces mostly to the operations and liquidity sections, exactly the areas where a tone-fundamentals gap would be expected to concentrate if the underlying business were softening faster than management wanted to say directly.
The value of a forensic disclosure score isn't the number itself. It's whether an analyst, a proxy advisor, or a journalist can trace that number back to the specific sentences, filings, and comparisons that produced it, and defend that trace in a governance conversation or a published story.
Readers building out their own workflow can find more detail on the underlying methodology in Lacuna's write-up on narrative engineering in corporate reporting and in the practical walkthrough of engineered disclosure patterns.
Building a Monitoring Pipeline Analysts Can Actually Run
A workable pipeline doesn't require a data science team, but it does require discipline about sequencing and documentation.
Start with the data sources: 10-K risk-factor sections, MD&A, earnings-call transcripts, and press releases, pulled for the current filing and at least three prior comparable periods for the same issuer.
- Run readability scoring first, average sentence length, hedge-term density per hundred words, and heading specificity against an ERM-style taxonomy.
- Layer in hedge counts by category (legal, operational, market, regulatory) so a spike can be traced to a specific risk domain rather than the filing as a whole.
- Generate sentence embeddings for each risk-factor paragraph and compute semantic distance against the issuer's own rolling baseline and its peer set.
- Feed the combined feature set into an anomaly detector and flag anything above a set percentile threshold.
- Decompose any flagged anomaly with a feature-level explainer and write the result into an evidence memo before escalating.
Reasonable starting thresholds: average sentence length above recommended limits, semantic-distance scores significantly higher than the issuer's recent baseline, and any instance of new generic risk language appearing alongside high overall textual similarity to the prior year's filing.
Pro Tip: Document the baseline period you used for every flag. When a governance committee or an editor asks why a filing was flagged, "compared against its own trailing three years" is a far stronger answer than "the model said so."
The evidence memo itself should include the specific sentences or sections flagged, the quantitative scores behind the flag, the corroborating checks run (or explicitly not run), and a recommended next action: escalate for deeper review, monitor next quarter, or close as a false positive. Analysts working across a broader coverage list benefit from standardizing this memo format using a shared disclosure-inconsistency toolkit so results are comparable across companies and reviewers.
Before-and-After: What Dilution Looks Like on the Page
The clearest way to recognize dilution is to see it happen across two filing cycles. Consider a hypothetical but representative pair of risk-factor excerpts.
A diluted version of the same risk, one cycle later, often reads: "Disruptions at our manufacturing facilities could adversely affect our results of operations." The location is gone. The concentration figure is gone. The impact range is gone. The hedge word "could" survives in both versions, so a pure sentiment count would show no meaningful change, while a human reader, or a semantic-distance model, sees a real loss of specificity.
Heading language shows the same pattern. "Risks Related to Our Dependence on [Named Supplier]" becoming "Risks Related to Our Supply Chain" over successive filings often signals that a named, trackable exposure has been folded into a generic catch-all, even when the underlying business relationship hasn't changed. The SEC's own EDGAR archive offers plenty of real examples of how risk-factor language and heading structure evolve across filing cycles, useful reference points for building an internal library of before-and-after comparisons specific to a coverage sector.
Why Dilution Happens: The Usual Culprits
Litigation defense is the most persistent driver. Legal teams reviewing draft risk factors tend to favor broader, vaguer language because a specific, quantified risk statement creates a more concrete target for a securities suit if the risk materializes. Regulators have started pushing back on this incentive directly; recent commentary on SEC efforts to trim corporate risk disclosures frames much of the current boilerplate expansion as filings functioning as litigation shields rather than investor communication.
Template reuse compounds the legal-defense incentive. Once outside counsel drafts a risk-factor template for one filing, it tends to get copied forward with minor edits for years, regardless of whether the underlying business risk has changed shape or size.
Generative-AI drafting tools have accelerated both problems. A model trained to produce compliant-sounding risk language will default toward the safest, most generic phrasing available, and the 2024 narrative-manipulation framework specifically flags AI-assisted drafting as a factor in the recent acceleration of boilerplate drift across filers.
Turnover on the disclosure-drafting side matters too. When general counsel, investor relations leads, or outside securities counsel change, institutional memory about why a specific risk factor was worded a particular way often leaves with them, and successor teams default to safer, vaguer replacements.
Finally, sheer volume plays a role. Risk-factor sections have grown long enough at many large filers that maintaining specificity across dozens of risk categories every year becomes a genuine drafting burden, one that generic language solves at the expense of investor clarity.
Measuring How Much Dilution Has Accumulated Over Time
A single filing snapshot tells you whether language is generic. A time series tells you whether it's getting worse, and by how much.
The most direct technique tracks semantic distance release over release: compute the embedding-based distance between each year's risk-factor section and a fixed early baseline (typically three to five years back), then plot the trend. A steadily rising distance score, even a modest one each year, compounds into a substantial narrative shift over a multi-year window that a single year-over-year comparison would understate.
Textual similarity scores run in the opposite direction and matter just as much. Rising similarity to the prior year's filing, paired with falling similarity to the company's own multi-year baseline, indicates the language has locked into a static template rather than evolving with the business, exactly the boilerplate-drift pattern worth tracking independently from semantic dispersion.
Hedge-density trend lines add a third dimension. Charting the ratio of hedge terms per hundred words across five or more filing cycles, broken out by risk category, shows whether hedging is increasing broadly or concentrated in one area, such as liquidity or regulatory risk, that deserves closer attention.

Combining these three time series, semantic distance from baseline, similarity to prior year, and hedge density by category, gives analysts a composite dilution trend rather than a single point-in-time judgment, and it's the structure Lacuna Index uses when comparing execution scores across a company's own multi-year history.
When Diluted Disclosures Become a Compliance Problem
Risk wording dilution sits close to, but is distinct from, outright disclosure fraud. The legal exposure it creates depends heavily on intent and materiality, but the pattern itself carries real regulatory weight.
Regulators evaluating whether a disclosure met its statutory purpose look at whether a reasonable investor would have understood the actual risk from the language provided. A risk factor stripped of its named specifics, its dollar exposure, and its quantified impact range can still be technically true while failing that reasonable-investor standard, which is precisely the gap that has drawn recent regulatory attention. Commentary from a former SEC chair has directly called out overlong, defensively drafted risk-factor sections as functioning more as litigation shields than investor communication tools, and that framing signals where enforcement priorities may shift.
The accounting-quality link adds a second layer of exposure. Since less-readable risk disclosure correlates with more aggressive discretionary accruals, a pattern of dilution can serve as a leading indicator that invites closer scrutiny of a company's broader financial reporting, not just its narrative language.
For governance professionals, dilution also raises board-oversight questions. A board that approves a proxy statement or 10-K without any process for checking risk-factor language against operational reality is exposed to the argument that it failed a basic oversight function, independent of whether any individual disclosure crosses into actionable misrepresentation.
Bringing Detection Into an Existing Review Workflow
The methods described above only create value if they sit inside a workflow someone actually runs on a schedule, rather than as a one-off research exercise.
The most sustainable integration point is the existing filing-review calendar most institutional analysts and governance teams already maintain. Running the readability and hedge-density checks as a first pass immediately after each 10-K or proxy filing hits EDGAR keeps the workload light and catches obvious cases within days rather than months.
The heavier semantic-embedding and anomaly-detection layer fits better as a quarterly batch process across a full coverage list rather than a per-filing task, since building and maintaining stable multi-year baselines for each issuer is where most of the computational and data-management effort lives. Running that batch process on a fixed calendar, rather than only when something already looks suspicious, is what catches slow-moving dilution before it becomes an obvious headline risk.
Documentation discipline matters as much as the technical pipeline. Every flag, whether generated by a simple readability threshold or a full anomaly-detection score, should land in a standardized evidence memo before it reaches a portfolio manager, an editor, or a governance committee. Journalists and governance advocates working across multiple companies benefit from the same standardized approach that forensic disclosure toolkits recommend: consistent thresholds, consistent baseline periods, and a consistent format for presenting the evidence trail behind any published claim.
Detection in Practice: What a Caught Case Looks Like
The clearest demonstrations of this methodology come from tracking a single issuer's risk-factor language across a multi-year window rather than a single filing. In a representative pattern, a company's semantic-distance score against its own three-year baseline rises steadily for two consecutive filing cycles while its hedge density in the liquidity risk category climbs in parallel, all while earnings-call transcripts maintain a consistently confident tone. The anomaly detector flags the filing well before any headline event, restatement, covenant breach, or credit-rating action, occurs.

The corroborating check in that pattern typically comes from outside the risk-factor text itself: a widening gap between reported earnings and cash flow from operations, or a subtle increase in days-sales-outstanding, that lines up with the timing of the textual anomaly. Decomposed with a feature-level explainer, the anomaly traces cleanly to a small number of sentences, not the entire filing, which is what makes the flag defensible in a governance conversation rather than a vague accusation of poor disclosure quality.
The mitigation side of these cases tends to play out through engagement rather than enforcement in most instances: a governance advocate or proxy advisor raises the specific flagged language directly with the company's investor-relations team, citing the exact sentences and the baseline comparison, and a subsequent filing cycle restores some of the lost specificity. That outcome, restored specificity following direct, evidence-backed engagement, is the practical goal of this entire detection exercise, and it depends entirely on the flag being explainable rather than a black-box score nobody can defend in the room.
Turn Detection Into a Standing Practice, Not a One-Off Audit
Spotting one diluted risk factor in one filing is a useful exercise. Building the infrastructure to catch dilution as it accumulates across a full coverage universe, quarter after quarter, is what actually changes an investor's or a journalist's information advantage. That requires baseline data most individual analysts don't have time to assemble manually: multi-year semantic histories, sector-level comparisons, and pre-scored narrative-versus-delivery gaps across hundreds of filers.
Lacuna Index's sector benchmarks provide exactly that starting point, a public reference for how narrative-delivery gaps distribute across peer groups, before any individual company report is pulled. Institutional subscribers gain access to the full company-level forensic reports, including the execution scores and audit-traceable evidence trails referenced throughout this piece. Readers new to the platform can start with the user guide on interpreting a Lacuna report to see exactly how a scored anomaly connects back to the specific filing language behind it.
Why Explainability Beats Alerts
Generative-AI drafting tools and rising litigation incentives are pushing risk disclosure toward more boilerplate, not less, and that trend will likely accelerate before regulators catch up to it. The temptation in response is to build louder alert systems. That's the wrong instinct.
An anomaly score nobody can trace back to specific sentences is worthless in a governance meeting or a published investigation. Explainable decomposition and stable, multi-year baselines aren't a nice-to-have layer on top of detection, they're what makes a flag defensible rather than merely suspicious. Analysts serious about this work should build their own methodology literacy through resources like Lacuna's narrative-engineering research and the user guide before trusting any single score.
— Glen
Sources
- A 2024 framework for detecting narrative manipulation in U.S. 10-K filings (arXiv preprint)
- Modeling structural deviation in 10‑K risk factors: a semantic anomaly detection and explainable AI approach (2026 preprint)
- SEC risk factors disclosure analysis (Harvard Law School Forum on Corporate Governance)
