AI Meeting Notes Evaluation Protocol: A Reproducible 12-Case Test
An AI meeting-note product should be evaluated as a record-making system, not as a writing demo. A summary can sound polished while changing a decision, dropping an owner, converting a tentative idea into a commitment, or sharing notes with people who should not receive them. The protocol below turns those risks into 12 repeatable cases and a score that another reviewer can reproduce.
The test is intentionally vendor-neutral. It works for a feature inside a meeting platform, a bot that joins calls, or a service that processes recordings later. Use synthetic meetings or recordings created for evaluation. Do not upload a real confidential call merely to test a product.
Set the acceptance rule before the demo
Write the decision rule before anyone sees product output. Otherwise, a fluent summary can move the goalposts. Start with four hard gates: participants are notified; the approved account and storage region are used; access and retention are configured; and a human owns final distribution. A product fails the pilot if any hard gate fails, even if its prose score is high.
Then define numerical thresholds. A reasonable internal pilot might require zero wrong decisions, zero invented action items, at least 95 percent capture of critical names and numbers, and no more than 10 minutes of correction per meeting. Those are examples, not universal standards. Legal, medical, safety, financial, or employment meetings usually need stricter review and may be unsuitable for automated notes.
Build a 12-case evaluation pack
Create four short scripted meetings and run each under three audio conditions. Keep the script, audio, participant labels, product settings, output, and human corrections together. The four content cases should be:
- Decision case: two options are discussed, one is rejected, and one is approved with a reason.
- Action case: three tasks have different owners, dates, and dependencies; one suggestion is explicitly not assigned.
- Entity case: include product names, acronyms, currency, percentages, dates, and similar-sounding names.
- Sensitive case: announce an off-record segment, stop note capture if the product permits it, then resume with a harmless topic.
For each script, produce a clean single-speaker recording, a normal remote-call version with interruptions, and a difficult version with background noise or an accented speaker. That gives 12 cases without pretending one tidy sample represents production.
Create the human reference correctly
Two reviewers should produce the reference independently from the audio, then resolve disagreements. The reference is not a verbatim transcript alone. It needs structured fields: decisions, rejected proposals, action text, owner, due date, unresolved questions, critical entities, and distribution restrictions. Preserve uncertainty. If the group said “probably Friday,” the reference must not silently change it to a firm Friday deadline.
Blind the reviewer who scores product output to the vendor name when practical. Remove branding and randomize output order. This reduces the chance that familiarity or price influences the factual score.
Use weighted errors instead of one vague quality score
| Dimension | What to count | Suggested weight |
|---|---|---|
| Decision fidelity | Correct decision, rejected options, conditions, and uncertainty | 30% |
| Action fidelity | Task, owner, date, dependency, and no invented assignment | 25% |
| Critical entities | Names, amounts, percentages, identifiers, and dates | 20% |
| Coverage | Important topics present without irrelevant padding | 10% |
| Speaker attribution | Statements assigned to the correct participant | 5% |
| Edit effort | Minutes from raw output to approved notes | 10% |
Mark a wrong decision or invented commitment as a critical error, not a minor wording defect. Keep a separate count for omissions and fabrications. A product that omits two low-priority comments is materially different from one that invents an approval.
Test consent, access, and retention as product behavior
Documentation is the starting point, but the configured workflow must also be observed. Google documents that its meeting-note feature can require participant consent, stores the generated document in the organizer's Drive, follows the organization's Meet retention policy, and currently supports one meeting language at a time. Microsoft documents different behavior depending on whether transcription is enabled: in-meeting assistance can be available without recording or transcription, while post-meeting access relies on transcript-related data. These differences belong in the test log because “AI notes” is not one uniform data flow.
Record who can start and stop capture, who receives the result by default, whether external invitees can see an attachment, where the file lands, how an administrator deletes it, and whether backups or derived summaries follow the same retention rule. Otter documents a configurable retention policy; confirm that the control exists on the exact plan under evaluation rather than assuming a help article applies to every account.
Run the pilot in a fixed order
- Freeze product version, plan, language, settings, date, and account type.
- Run the 12 recordings without changing prompts between vendors.
- Export raw notes immediately and preserve timestamps where available.
- Have two reviewers score factual fields before editing style.
- Measure correction time with a start and stop timestamp.
- Repeat any failed case once to distinguish a persistent failure from a transient service error.
- Document the go, limited-go, or no-go decision and its expiration date.
A limited-go outcome is often more honest than a universal approval. For example, a system may be allowed for routine internal project calls but prohibited for interviews, disciplinary discussions, regulated data, customer secrets, or multilingual meetings.
Report failure patterns, not only averages
Tag every defect as omission, fabrication, attribution error, entity error, certainty inflation, privacy-control failure, or service failure. Then report counts by content case, audio condition, and language. An 88 percent overall score can conceal that every noisy call loses due dates or that one speaker is consistently assigned another person's actions. Include representative before-and-after corrections with synthetic details so stakeholders can see what the number means.
The final report should name the approved scope, human reviewer, prohibited meeting types, fallback when capture fails, correction workflow, retention owner, and next review date. Approval should expire after a material feature, model, plan, sharing-default, or retention change. That turns a one-time demo into a controlled service decision.
Original worksheet
Download the AI meeting-notes evaluation sheet (CSV). It includes all 12 case rows, critical-error fields, correction time, consent, access, retention, reviewer, and decision columns. Duplicate the blank rows for additional languages or audio conditions. Keep raw outputs outside the CSV and link them with a non-sensitive test ID.
Source map
- Google Meet Help supports the documented language, consent, sharing, storage, and retention behavior cited above.
- Microsoft's Teams meeting guide and its no-transcription guide support the distinction between in-meeting and post-meeting use.
- Otter's retention article supports the existence of administrator-configured retention on applicable accounts.
- NIST AI RMF supports the risk-based govern, map, measure, and manage framing used for the protocol.
Limitations
This protocol was source-reviewed on July 17, 2026; no vendor accuracy benchmark was performed for this article. Product plans, interfaces, languages, data locations, and retention controls can change. A 12-case pilot detects obvious workflow failures but cannot represent every speaker, disability, jurisdiction, meeting type, or adversarial condition. Consent and recording law vary by location, and this article is not legal advice. Obtain legal, security, accessibility, and records-management review for the intended use.
Primary sources checked
- Google Meet Help: Take notes for me
- Microsoft Support: Use Copilot in Teams meetings
- Microsoft Support: Use Copilot without transcription or recording
- Otter Help: Set a custom data-retention policy
- NIST AI Risk Management Framework
Product details and prices can change after the review date. Verify the linked official page before purchasing or deploying.