meeting operations17 min read

Nobody said September: the nine-minute review every AI meeting summary needs

Research on meeting-transcript summarization found that 30.4% of generated summaries of city-council meetings carried at least one factual error, and seven distinct error types behind them. A practical guide to reviewing an AI meeting note against its own transcript: what to check, in what order, and what to say in the room so the review takes nine minutes instead of thirty.

K
Ken Jo
#meeting-notes#ai-notetaker#ai-summary#meeting-minutes#note-taking#hallucination#transcription

The summary was three paragraphs long, cleanly organized, and wrong in exactly one sentence: the team agreed to ship the migration in September.

Nobody had agreed to that. Someone had said that September was probably doable if the vendor came back in time — a conditional, spoken once, in the middle of a tangent about something else. The model dropped the condition, promoted the guess to a decision, and the note went to eleven people who weren't in the room. Two of them started planning against September.

Nothing about that is exotic. It is one of the most thoroughly catalogued behaviours of automatic summarizers, it has a name in the research literature, and it is not going away because the next model is bigger. This post gives you three things: the seven ways a summary reliably drifts from what was actually said, a nine-minute review that catches all seven, and the handful of sentences you can say during a meeting that make the review shorter.

TL;DR:

  • In the TofuEval benchmark (NAACL 2024), 30.4% of generated summaries of real city-council meeting transcripts contained at least one factual error, and for the marginal topics — the tangents — that rose to 43.6%.
  • Errors are not random noise. Researchers sorted them into seven named types, and the two that hurt meetings most are stating opinion as fact and tense/modality shifts — exactly how "we might" becomes "we will."
  • You cannot hand the checking back to a model: on FaithBench (October 2024) the best hallucination detectors scored near 50% accuracy, and in TofuEval, LLM evaluators were beaten by smaller specialized metrics.

An open notebook with several pens and a pair of reading glasses resting on it, on a desk lit for evening work

Image: Shixart1985, Wikimedia Commons, CC BY 2.0. The glasses are the point. Somebody still has to read the thing.

Capture got solved. Checking never got started.

Search Hacker News for meeting-note tools launched in the last six months and you can count at least fifteen between March 1 and August 13, 2026 — on-device transcribers, bot-free notepads, Obsidian importers, open-source Granola clones. Every one of them ships the same two features: capture the audio, generate the summary.

Not one of them ships the third step, which is the only one with any judgment in it. The summary lands in your inbox already formatted, already confident, already sent to the room. The implicit claim in that workflow is that the output is finished.

Here's the point: a generated summary is not a record of the meeting. It is a statement about the meeting, produced by a system that was not accountable for being right, and it should be read the way you'd read any other secondhand account — against the evidence it came from.

The evidence, thankfully, is right there. Every tool that produced the summary also produced the transcript, and reviewing one against the other takes less time than the meeting spent on introductions.

The number that should change how you read your notes

TofuEval, published by researchers at Salesforce AI and collaborators at NAACL 2024, is the closest thing we have to a controlled measurement of this. They took two dialogue corpora — one of news interviews, one called MeetingBank made of real US city-council meeting transcripts — had five language models write topic-focused summaries, and had professional linguists annotate every summary sentence for factual consistency against the source.

On the meeting transcripts, averaged across the five models, 30.4% of summaries about the meeting's main topics contained at least one factual inconsistency, rising to 43.6% for marginal topics. The paper states it plainly: apart from GPT-3.5-Turbo, "approximately 40–50% of their summaries contain at least one factual inconsistency."

The honest caveat: those five models are from 2023 — Vicuna, three sizes of WizardLM, and GPT-3.5-Turbo — and today's frontier models are meaningfully better at faithfulness. Two findings from the same paper survive that objection, though, and they're the ones that matter for your Tuesday.

First, scale did not reliably fix it. WizardLM-30B's error rate on the meeting data was only 3.8 percentage points better than WizardLM-7B's, despite being more than four times the size, and on some comparisons the larger model in a family produced more errors. Second, error count had no meaningful relationship to summary length (Pearson ρ = 0.18) — a longer, more detailed summary is not a safer one. Waiting for a bigger model is not a review process.

Bar chart: share of generated meeting summaries containing at least one factual error, by model, ranging from 10.9% to 41.3%

Summaries of real city-council meeting transcripts containing at least one factual error, main topics only. Source: TofuEval, Tang et al., NAACL 2024 (arXiv:2402.13249). Models tested are 2023-era; the error types below outlived them.

Seven ways a summary drifts from the room

The most useful thing in that paper is not the percentage. It is the taxonomy — seven named error types the annotators derived from thousands of flagged sentences, extending an earlier taxonomy built on short dialogues.

Read them once and you will start seeing them in your own notes, because they are not random glitches. Each one is a specific, repeatable way that compressing speech into prose loses a distinction the room was relying on.

Error type (TofuEval)What it looks like in a meeting noteHow you catch it
Extrinsic informationA detail that appears in the summary and nowhere in the audio — a number, a customer name, a rationale nobody gaveSearch the transcript for the specific noun or figure; if it returns nothing, it was invented
MisreferencingThe right decision attributed to the wrong person, or one person's objection credited to anotherJump to the timestamp and read who is speaking
Stating opinion as fact"The pricing change will increase churn" — when someone worried it mightLook for the hedge in the original: thought, worried, suspects
Reasoning errorWrong arithmetic, or a cause-and-effect link the room never drewRecompute any number; ask whether the "because" was said or inferred
Tense / aspect / modality"We could ship in September" becomes "we will ship in September"Check every future-tense verb against the original modal
ContradictionA dropped negation: "we are not moving the deadline" becomes "we are moving the deadline"Grep the transcript for not, never, unless near the claim
Nuanced meaning shift"Make a recommendation" paraphrased as "make a request"Compare the verb in the summary to the verb in the source

Three of them are easiest to recognize as artifacts. Here is each one with the transcript and the summary line placed side by side.

A dropped negation — contradiction error. What the room said:

DANIEL (00:31:12): So we are not moving the launch date. Whatever
else changes, that date holds.

What the summary said:

The team agreed to move the launch date.

One word gone, and the sentence now instructs the opposite action. It is also grammatically perfect, which is exactly why proofreading sails past it while a search for "not" near the decision catches it in four seconds.

A misattribution — misreferencing error. What the room said:

PRIYA (00:18:40): I'd push back on per-seat. Two enterprise buyers
told us last quarter that it kills the deal.
MINA (00:18:55): Yeah, agreed.

What the summary said:

Mina objected to per-seat pricing, citing enterprise buyer feedback.

Every fact in that sentence is in the transcript. The person is not. That is the whole difficulty: you cannot catch this by asking whether the claims are true, only by asking whether they belong to the person they were pinned on.

A promoted guess — modality error. What the room said:

SAM (00:44:02): September's probably doable, if the vendor comes
back to us in time.

What the summary said:

The team agreed to ship the migration in September.

Two qualifiers removed — the word "probably" and the entire conditional clause — and a shared guess became a commitment with a date attached. Nobody had to lie for that to happen. The compression did it.

Two of these deserve special attention, because they are the ones that cost teams real money.

Stating opinion as fact is the September failure from the top of this post. In a meeting, most sentences are hedged — people are thinking out loud, and the hedge carries the whole meaning. Summarizers strip hedges by design, because hedges are the first thing you cut when compressing prose. The result reads more decisive than the room was, which is precisely why it goes unquestioned.

Modality shifts do the same damage to dates. "Could," "should," "probably," and "if the vendor gets back to us" are the difference between a plan and a commitment, and they are also the cheapest words to drop. A summary that says the team will do something in September has quietly manufactured a deadline nobody owns.

Notice what both have in common. The model did not invent facts out of nothing. It removed qualifiers — the least visible edit it can make, and the one no reader can detect without the source in front of them.

The tangents are where it invents things

There is a second pattern in the data that maps directly onto how meetings actually run.

Summaries of marginal topics — the things mentioned in passing rather than discussed at length — were substantially worse: 43.6% versus 30.4% on the meeting corpus. The researchers explain the mechanism: when a topic is barely covered in the source, models "rely on their knowledge to make informed inferences about the topic, bringing unsupported information into the summary." Thin evidence gets padded out from general knowledge, and the padding is indistinguishable in tone from the parts that are grounded.

Your meetings are full of marginal topics. The two-minute aside about the security review, the vendor mentioned once, the number somebody half-remembered. Those lines are the least discussed and the most likely to be wrong, and they will read exactly as fluently as the rest.

The layer below has the same shape. In "Careless Whisper," presented at ACM FAccT 2024, Koenecke and colleagues found that roughly 1% of Whisper transcriptions contained entire hallucinated phrases or sentences that did not exist in any form in the underlying audio, and that 38% of those hallucinations carried explicit harms — invented associations, false authority, fabricated violence. The trigger was the interesting part: hallucinations occurred disproportionately for speakers with longer stretches of non-vocal time.

Silence, in other words. Which is what a meeting is made of — pauses to think, someone unmuting, a screen share loading, the four seconds after a hard question. That is the raw material a transcription model is most likely to fill in with something that was never said. It also means transcript errors and summary errors are not independent: an invented transcript line is a perfectly grounded summary input.

You cannot hand the checking back to the model

The obvious idea is to have a second model verify the first one. The measurements are not encouraging.

In the same TofuEval study, the authors evaluated LLMs as binary factual-consistency judges and found that all of them, GPT-4 included, "perform poorly at detecting errors in LLM-generated summaries" — and were outperformed by smaller, non-LLM factuality metrics built specifically for the task. Even when GPT-4 correctly flagged a sentence, its explanation of why was right about 80% of the time; for the other evaluators, roughly half.

FaithBench, released by Vectara in October 2024, put a number on the ceiling. It collected the hallucinations that ten modern LLMs across eight model families produced and that existing detectors disagreed about — and reported that even the best hallucination detection models scored near 50% accuracy on them. A coin flip, on the cases that matter.

So the review is yours. That sounds like a burden until you notice the asymmetry: the detector has to judge arbitrary text against an unfamiliar document, while you were in the room. You already know which claim looked odd. You only need to check it.

Three artifacts, three different jobs

Most "AI notes versus real notes" arguments are a category error. A meeting produces three artifacts, they are not substitutes, and a team that keeps only one always loses something specific.

Raw transcriptAI summaryDecision log
Exact wording preservedYes — verbatimNo — compressed, hedges droppedNo
Who said itYes, if diarizedSometimes; misreferencing is a named error typeYes — owner is the point
When it was saidYes — timestampedRarelyOnly the deadline
Decisions with an owner and a dateBuried in ~9,000 wordsPartially, and unreliablyYes — this is its only job
What was ruled out and whyYes, if anyone said itUsually dropped as redundantYes, in one line
Readable in 60 secondsNoYesYes
Can be checked against evidenceIt is the evidenceOnly if the transcript survivesOnly if both survive

Matrix: transcript, AI summary and decision log scored on six properties, showing that no single artifact covers all of them

What each artifact keeps, partially keeps, or loses. The bottom row is the one that decides whether the other two are worth anything.

The bottom row is the whole argument. A summary with no surviving transcript is unfalsifiable — when a colleague says "that's not what we agreed," you have two opinions and no record. This is where an AI note is more fragile than a human one, not less: a human note-taker's minutes came with visible fallibility attached, and everyone read them with appropriate suspicion. Machine prose arrives in clean headings and confident declaratives, and borrows authority it did not earn.

Call it what it is: the summary is a witness, and the transcript is the evidence. Testimony is useful. Testimony is also the thing you check.

The nine-minute review, in order

Here is the procedure we recommend, timed for a 60-minute meeting. The order matters — it front-loads the checks with the highest cost of being wrong, so that if you get interrupted at minute four you have already caught the expensive errors.

  1. 0:00–1:30 — Read only the decision sentences. Ignore the narrative summary entirely on the first pass. Find every sentence that asserts an outcome, a commitment, or a date. In a typical 60-minute meeting there are three to six. These are the only lines anyone will act on.
  2. 1:30–3:00 — Restore the modality. For each of those sentences, jump to its point in the transcript and check the verb. Did someone say will, or did they say could, should, or probably? If the original was hedged, rewrite the summary line with the hedge back in — or better, mark it as an open question with an owner.
  3. 3:00–4:30 — Check every proper noun and number. Names, companies, versions, prices, dates, percentages. Search the transcript for each one. Anything that does not appear there is extrinsic information: delete it or confirm it out of band. This is a search-box task, not a reading task.
  4. 4:30–5:30 — Verify attribution. For each decision and each objection, confirm who is speaking at that timestamp. Misreferencing is the error most likely to cause an interpersonal problem rather than a planning one, and it is the fastest to check.
  5. 5:30–6:30 — Hunt for dropped negations. Search the transcript for not, never, don't, unless, and instead of in the neighbourhood of each decision. A contradiction error inverts the meaning while leaving the sentence perfectly grammatical, so proofreading will not find it — searching will.
  6. 6:30–8:00 — Ask what is missing. This is the step no tool can do for you, because a summary cannot flag its own omissions. Two questions: what did we rule out, and why? And what question did we fail to answer? Neither survives compression, and both are what makes the note useful in six months.
  7. 8:00–9:00 — Write the decision log and send it. Three to six lines, each carrying what was decided, who owns the next move, and by when. Link it to the transcript. That link is what makes the whole review re-checkable by somebody else.

Timeline: a nine-minute review split into seven steps, from checking decision sentences to sending the decision log

The order is the design: modality and invented specifics first, omissions last. An interrupted review still catches the expensive errors.

Two notes on doing this at scale. If the meeting was 30 minutes, this is a four-minute job. And if you do it in the tool that holds both the summary and the timestamped transcript, steps 2 through 5 are all searches and clicks rather than window-switching — which is the difference between a habit that survives a busy week and one that doesn't.

Ten seconds in the meeting saves five minutes in the review

The review gets dramatically shorter if the source recording contains the information in the first place. A model cannot recover what nobody said, and most of what makes notes ambiguous was never in the audio.

  1. Open with the question, out loud. "We're here to decide whether the migration ships before or after the conference." A stated question gives the summarizer a main topic instead of a set of marginal ones — and marginal topics were the 43.6% case.
  2. Say the decision in one sentence, with a name and a date, before anyone leaves. "Decision: Mina owns the vendor follow-up, and we confirm the September date by the 25th." Every notetaker on the market captures that perfectly. Not one will invent it.
  3. Say the negation explicitly. "We are not moving the launch date" is safer than a shrug and a nod, because a dropped negation is a documented error type, and an unspoken one cannot be dropped — it was never there.
  4. Name what you rejected and why, in one line. "We considered per-seat pricing and rejected it because two enterprise buyers pushed back last quarter." This is the sentence that stops the same debate reopening in six months.
  5. Spell unusual proper nouns once. Vendor names, internal codenames, and acronyms are exactly what a transcript garbles and a summary then confidently repeats.
  6. Don't let silence do the talking. If the room goes quiet while everyone reads a document, say so. Long non-vocal stretches are the conditions under which transcription models hallucinate most.

None of this is meeting theatre. It is dictating to the record — a skill that got more valuable when the record started being written automatically, not less.

The bottom line

An AI meeting summary is not a record. It is testimony about a record, delivered fluently, by a witness with no stake in being accurate and a documented tendency to drop the word "probably."

That reframing settles the manual-versus-AI argument, which was never a real argument. Keep the recording, because it is the only evidence. Let the model write the summary, because it is faster than you and it never gets bored at minute forty. Then spend nine minutes doing the one part that requires having been there: checking the confident sentences against what was actually said, and writing down the three lines that someone has to act on.

The teams that get burned in 2026 are not the ones that skipped the AI notetaker. They are the ones that read its output as a transcript when it was a paraphrase — and never kept the transcript that would have shown the difference. For the fuller version of what a good note has to carry, we wrote about the six things a meeting note still needs, and about why word error rate tells you almost nothing about whether a transcript is usable.


Where Telli.sh fits: the nine-minute review only works if the summary and the source live in the same place. Telli.sh keeps the timestamped transcript with speaker attribution next to an editable AI summary, so checking a claim is a search and a click rather than a hunt through two systems — and when the room isn't speaking one language, the translation sits alongside the original rather than replacing it. The summary arrives as a draft you correct, not an output you accept.

Start a live AI note

Sources


Back to Blog