← Back to blog

Earnings Call Credibility: A Forensic Audit Framework

August 26, 2026
Earnings Call Credibility: A Forensic Audit Framework

Earnings call credibility is a forensic, audit-traceable measure of whether management's spoken narrative matches its filed disclosures and delivered results, established solely from public records. The core empirical claim behind this measure is straightforward: divergence between what executives say on a call and what the filings actually show predicts post-disclosure outcomes, including returns. The immediate action for any analyst or governance professional is a baseline divergence check, comparing the call transcript against the MD&A section of the corresponding 10-Q or 10-K, and documenting the gap with quotes and timestamps.

That first pass should establish three things before any deeper scoring begins:

  • Whether management's tone on the call is measurably more optimistic than the language used in the filed MD&A discussing the same quarter.
  • Whether specific claims made verbally (margin trajectory, cost discipline, demand signals) appear, unqualified, in the filed numbers.
  • Whether guidance language on the call matches the risk factors and forward-looking disclaimers filed in the same period.

Key Takeaways

Earnings call credibility functions as a measurable, audit-traceable gap between spoken claims and filed evidence, not a subjective read of management's confidence.

PointDetails
Run the baseline check firstCompare the call transcript against the MD&A section for tone, complexity, and specific claims.
Segment prepared remarks from Q&AThe two segments carry different signal content, and Q&A often exposes hedging that prepared remarks conceal.
Require named assuranceRoughly 23% of firms claiming assurance name no provider, so verify provider, standard, and scope before trusting the claim.
Corroborate before actingSingle-call divergence is noise; validate across multiple calls and confirm against later filings, guidance cuts, or restatements.
Use Lacunaindex for scaleLacunaindex converts this protocol into execution scores, archetype classifications, and free sector benchmarks for institutional review.

Table of Contents

Why Earnings Call Trustworthiness Is a Measurable, Not Rhetorical, Question

The case for treating financial call reliability as a quantifiable construct rests on a growing body of matched-disclosure research, not intuition about who sounds confident on a call. A study analyzing 22,366 matched disclosure pairs across 1,762 firms from 2006 to 2025 found that divergence in tone and complexity between earnings calls and MD&A filings predicts post-disclosure stock returns, with complexity gaps producing effects that persist well beyond the initial reaction window.

Statistic Callout: Across more than 22,000 matched pairs of call transcripts and MD&A filings, tone and complexity divergence between the two documents was a significant predictor of post-disclosure returns, not just short-term noise.

A separate validation effort behind the Disclosure Authenticity Evaluation Model (DAEM) tested whether communication authenticity, meaning operational alignment, temporal consistency, and specificity, could be measured reliably across raters. It could: the framework achieved an intraclass correlation of 0.85 and Krippendorff's alpha of 0.83 across eight mega-cap companies, a level of inter-rater agreement that puts credibility scoring on the same methodological footing as other validated behavioral measurement systems.

Two implications follow for institutional users:

  • Divergence signals are not purely a short-term trading input. The same DAEM research found authenticity correlated more strongly with employee engagement (r = 0.423) than with abnormal stock returns (r = 0.289, not statistically significant), meaning credibility gaps often show up first in stakeholder trust and only later, if at all, in price.
  • Reporting quality assessments built on these signals give governance professionals and proxy advisors a defensible, evidence-based basis for engagement, rather than a subjective read of management's tone on a call.

A Step-by-Step Protocol for Scoring Earnings Call Integrity

Evaluating earnings call trust at institutional scale requires a repeatable sequence, not an ad hoc read of the transcript. The protocol below uses only public records and produces evidence that can be checked line by line.

  1. Collect and align every artifact for the period. Pull the call transcript, the press release, the Item 2.02 or 8-K filing, the 10-Q or 10-K MD&A, and, where the calendar allows, the proxy statement. Timestamp each document and extract exact quotes rather than paraphrases.

  2. Segment the call before scoring anything. Prepared remarks and the Q&A session carry different signal content. Linguistic and vocalic analysis of earnings-call utterances has found systematic differences between restatement-related and nonrestatement utterances, and prepared remarks versus spontaneous answers differ meaningfully in what they reveal. Score them separately.

  3. Establish per-speaker historical baselines. A CFO who suddenly adopts new hedges, qualifiers, or evasive phrasing relative to their own prior calls is a more useful signal than a one-off comparison against peers. The Spoken Alpha methodology for behaviorally grounded deviation flags demonstrates how longitudinal, speaker-specific baselines catch departures that cross-sectional comparisons miss.

  4. Measure tone, complexity, and guidance divergence. Score sentiment gaps between the call and the MD&A, run readability comparisons, and check whether forward guidance stated verbally matches the risk language filed in the same period.

  5. Map narrative claims to hard metrics. Every specific claim, margin improvement, capex discipline, unit growth, needs a corresponding number in the filed statements. Text-mining approaches that pair qualitative tone with quantitative performance clusters can systematically classify whether a "positive" narrative sits on genuinely strong numbers or a weaker underlying result.

  6. Validate against confirmatory filings. Guidance cuts, pre-announcements, and restatements in subsequent quarters are ground-truth events. A credibility flag that later coincides with one of these carries far more weight than one that does not.

  7. Document everything. Every score needs a quote, a timestamp, and a direct link to the specific filing item used to confirm or refute the claim. Evidence packages built this way survive scrutiny from a compliance committee or an editor in a way a summary judgment never will.

Pro Tip: Score prepared remarks and Q&A as two separate line items in your workpaper, not one blended average. A CEO can deliver flawless prepared remarks and still contradict the filed MD&A the moment an analyst asks an unscripted follow-up question, and that gap is where most of the useful signal lives.

This structure mirrors what practitioners call dual-stream auditing, comparing qualitative tone against quantitative ratios while tracking the complexity gap between the two, an approach that reliably surfaces the kind of inconsistency that signals an engineered disclosure. For a deeper walkthrough of how narrative claims get tested against financial context, see this forensic corporate analysis primer.

What a Credibility Scorecard Should Actually Measure

A usable scorecard for evaluating earnings call trust needs a small number of measurable signals, each backed by a specific detection method, rather than a vague impression of "confidence."

  • Tone divergence. Sentiment shifts between prepared remarks and Q&A, and between the call overall and the filed MD&A, are the single most researched signal, and the one with the strongest link to post-disclosure returns.
  • Complexity gap. Readability differences, using measures like the Loughran–McDonald financial dictionary or standard Fog and Gunning indices, between spoken language and filed text often widen precisely when management wants a weak quarter to sound routine.
  • Per-speaker deviation. Novel hedges, unusual qualifier density, or a spike in hedge language relative to a speaker's own historical baseline flags a departure worth investigating.
  • Guidance consistency. Cross-check verbal guidance against the risk factors and forward-looking language filed in the corresponding 8-K or 10-Q.
  • Assurance naming. When management cites third-party assurance or a compliance standard, check whether a named provider, standard, and scope actually appear. Industry-wide review of disclosure practices found that 71% of firms claiming assurance provided one, yet 23.3% named no provider at all, a gap of 111 companies making a claim with nothing behind it.
  • Quantitative mismatches. KPIs cited verbally that don't reconcile with reported numbers, or unexplained one-off adjustments, are the clearest hard-metric red flag.

Statistic Callout: Nearly one in four companies claiming third-party assurance name no provider at all, turning a supposed credibility signal into an unverifiable assertion.

How Lacuna Index Turns These Methods Into Institutional Output

Lacunaindex operationalizes this protocol at scale by mining the same public-record inputs the audit framework calls for: call transcripts, SEC filings, press releases, and proxy statements. Each input feeds a specific check, transcripts for tone and per-speaker baselines, filings for hard-metric reconciliation, proxies for governance context, and every claim in the resulting report traces back to a quote, a timestamp, and a filing link.

The output takes three forms:

  • Execution and narrative scores that quantify the gap between what a company claims and what its filings confirm.
  • Archetype classification, sorting companies into categories such as earned, borrowed, or undervalued based on how execution compares to valuation.
  • Audit-traceable evidence packages, built so a governance committee, proxy advisor, or newsroom can verify every underlying data point.

Institutional users apply these outputs three ways: screening portfolios for narrative-versus-delivery gaps, prioritizing which companies warrant direct engagement, and, for journalists, sourcing accountability stories with evidence that holds up to scrutiny. Free sector benchmarks are available for readers who want a first look before subscribing to company-level reports.

OutputWhat it measures
Execution scoreGap between narrative claims and confirmed delivery
Archetype classificationEarned, borrowed, or undervalued positioning
Evidence packageQuote, timestamp, and filing link for each flagged claim

Common Tactics Companies Use to Obscure a Weak Quarter

Companies rarely fabricate numbers outright on a call; the more common move is to shape which numbers get emphasized and how. A frequent tactic is burying a weak segment inside an aggregate figure, citing consolidated revenue growth while a core division actually contracted, a pattern only visible when the call is checked against segment-level detail in the 10-K.

Another common pattern is guidance language that quietly shifts scope. A company might reaffirm "full-year targets" on the call while the filed risk factors already flag conditions that make the prior range unlikely, a divergence that shows up clearly once the transcript and the 8-K are read side by side.

Selective KPI citation is a third tactic: management highlights a metric that improved (units shipped, new logos, engagement) while omitting the metric that actually drives the P&L (average selling price, churn, gross margin) in the same breath. And complexity itself can be a tool. Loughran-McDonald based readability comparisons often show call language running noticeably simpler than filed MD&A language in strong quarters, and that gap narrowing or reversing when results disappoint, a pattern consistent with management reaching for denser, more hedged phrasing when the story is harder to tell plainly. None of these tactics require a false statement, which is exactly why a forensic, filing-based check catches what a purely rhetorical read of the call cannot. Readers who want more detail on how these patterns show up across sectors can review this narrative engineering analysis.

Common Tactics Companies Use to Obscure a Weak Quarter — overview diagram

Where Management Behavior and Vocal Cues Fit, and Where They Don't

Vocal cues and management demeanor generate real research interest, but they belong in a credibility assessment as a secondary signal, not a primary one. Linguistic and vocalic analysis of earnings-call utterances has found measurable differences between utterances later linked to restatements and those that were not, which means these markers carry genuine information, particularly when comparing the same speaker to their own historical pattern rather than to peers.

The practical distinction that matters for evaluating earnings call trust is prepared remarks versus Q&A. Prepared remarks are rehearsed, reviewed by counsel, and generally free of spontaneous hedging. Q&A is where analysts push, and where hedging, qualifier density, and evasive phrasing tend to surface if a claim doesn't hold up under a follow-up question. Treating vocal and behavioral cues as a standalone credibility verdict overstates what the research supports; treating them as one input inside a per-speaker deviation baseline, checked against the same speaker's prior calls, is where they earn their place. Behavior alone never substitutes for the reconciliation against filed numbers that anchors a credible score.

How Credibility Gaps Actually Move Investor Decisions

The practical payoff of scoring investor call credibility isn't a clean short-term trading signal, and institutional users should be skeptical of anyone who claims otherwise. The DAEM validation work found no statistically significant correlation between communication authenticity and abnormal stock returns in the event study it ran (r = 0.289, p = 0.491), while finding a meaningfully stronger association with employee engagement (r = 0.423). Credibility gaps tend to show up first in how stakeholders, employees, analysts, governance committees, respond, and only later, if the gap is severe enough, in price.

That timing gap is exactly why credibility scoring earns its keep for governance professionals and proxy advisors rather than day traders. A persistent tone or complexity divergence, corroborated across multiple quarters and confirmed against a later guidance cut or restatement, is a legitimate basis for an engagement letter, a voting recommendation, or a deeper diligence request, well before the market has fully repriced the stock. For journalists, the same evidence package supports accountability coverage that can withstand a fact-check, because every claim traces back to a specific quote and filing. The regulatory oversight gaps guide covers where these gaps most often escape scrutiny before they become material.

How Credibility Gaps Actually Move Investor Decisions — overview diagram

Limits and Common Failure Modes: An Advisor's Note

A single call flagged for divergence is noise, not a verdict. Macro shocks, accounting changes, and a new CFO's speaking style can all masquerade as a credibility problem. The fix is corroboration: require a named assurance scope before trusting any assurance claim, and validate flags across multiple calls or a portfolio, not one earnings season. Readers building engagement priorities should treat this as a companion practice to standard governance and transparency review, not a replacement for it. Escalating a confirmed pattern to formal investigation deserves its own rigor, which is where structured fraud investigation steps become relevant.

— Glen

Get Sector Benchmarks and Forensic Reports From Lacunaindex

The audit protocol above works, but running it manually across a full coverage list or watchlist takes real analyst hours that most institutions don't have to spare every quarter. Lacunaindex built its platform to do exactly this reconciliation at scale, mining transcripts, filings, press releases, and proxies to produce execution scores, archetype classifications, and audit-traceable evidence packages, so you get the output of the protocol without running each step by hand.

Lacunaindex

That fits governance professionals prioritizing engagement targets, proxy advisors building voting recommendations, and financial journalists who need a sourced evidence trail before a story runs. Start with the free sector benchmarks to see how companies in a given sector compare on execution versus narrative, then move to a subscription for company-level forensic reports when a specific name warrants deeper diligence.

Sources

Matched disclosure evidence, DAEM validation, and Spoken Alpha methodology are cited throughout the sections above.

This article is general information, not a substitute for advice from a qualified financial advisor. Consult a qualified financial professional about your own circumstances before acting on anything here.