The 60-second verdict
Quick answer: benchmark AI transcription accuracy with a locked, representative test set; a human-checked reference transcript; predefined scoring rules; and transparent denominators. Report word error rate, processing failures, critical-detail accuracy, speaker attribution, timestamps, unsupported additions, severity and condition-level results rather than one headline percentage.
Decision focus: use the method below only where it produces a recoverable source, a verifiable output and a clear next action. If one of those fails, change the workflow rather than trusting a polished summary.
Evidence basis and limits
- Decision factors covered: Define the decision and quality threshold; Lock and version the test set; Create a controlled reference transcript.
- Evidence rule: A claim earns weight only when the source, date, configuration and limitation are clear enough for a reader to check.
- Boundary: Examples and workflow recommendations must be tested with representative recordings, the intended users and the actual approval process before rollout.
A precise-looking accuracy score can still be misleading if failed files are excluded, easy audio dominates the sample or an incorrect amount is treated like harmless punctuation. The benchmark should support a specific operational decision and expose the errors that matter to that use.

Define the decision and quality threshold
State what the transcript will be used for and the quality required. A searchable personal draft may tolerate harmless wording differences; an approved meeting or professional record needs stronger controls for names, figures, decisions, conditions and speaker identity.
Set pass, conditional-pass and fail thresholds before testing. Include mandatory gates for processing reliability and high-risk fields instead of allowing a strong average to compensate for a serious defect.
Lock and version the test set
Create a manifest containing sample ID, duration, language, speaker count, room condition, microphone distance, recording mode, device, app version and specialist vocabulary. Include the real accents, languages and environments expected in use.
Write inclusion and exclusion rules in advance. Keep corrupt files, failed uploads, empty transcripts and incomplete outputs in the reconciliation record. Do not remove difficult samples after seeing the results.
Create a controlled reference transcript
Use a documented convention for filler words, false starts, punctuation, contractions, partial words, non-speech sounds and inaudible passages. A second reviewer should check material words, names, figures, negations and speaker changes.
Lock the reference version before scoring. If the reference later changes, recalculate the affected results and issue a new report version.
Calculate word error rate correctly
WER = substitutions + deletions + insertions ÷ reference words.
For example, 22 substitutions, 8 deletions and 5 insertions across 2,500 reference words produces a WER of 1.4%. Report all raw counts and the reference-word denominator.
WER can exceed 100% when there are many insertions. Do not present 100 minus WER as universal “accuracy” without explaining the limitation.
Measure processing failures
Processing-failure rate = failed files ÷ attempted files.
Define failure before testing. It may include an upload error, unsupported format, empty transcript, truncated output, unrecoverable timeout or application crash. Report successful reruns separately so reliability remains visible.
Reconcile every denominator
Keep a table showing recordings selected, attempted, processed and scored; exclusions; reference words; critical fields; failures; and reruns. If a denominator changes between reports, explain why and show the impact where practical.
Measure critical-detail accuracy
Create a locked list of names, organisations, dates, times, amounts, percentages, units, technical terms, identifiers, negations, decisions, action owners, deadlines and conditions.
Critical-detail accuracy = correct critical details ÷ total reference critical details.
Report correct, incorrect, omitted and unsupported fields. A field is not correct when the words look similar but the operational meaning changes.
Measure speaker attribution
Choose one scoring unit—speaking turn, sentence, timed segment or material statement—and use it consistently.
Speaker-attribution accuracy = correctly attributed units ÷ total reference units.
Report misattributed and unattributed units separately. A low WER does not protect against assigning a promise to the wrong person.
Measure timestamp accuracy
Set an acceptable tolerance before testing. Report the proportion of reference points within that tolerance and the median absolute timing error. Do not judge timestamps by eye after seeing the output.
Audit unsupported additions
For summaries or structured notes, trace every material claim to the source.
Unsupported-addition rate = unsupported material assertions ÷ total material assertions.
Define the assertion unit before review. Separate invented decisions, owners and reasons from ordinary word insertions.
Apply a fixed severity framework
- Cosmetic: harmless punctuation, filler or formatting.
- Operational: requires correction or weakens search and follow-up.
- High-risk: changes identity, amount, deadline, safety, obligation, decision status or another material outcome.
Publish raw counts alongside any weighted score. Weighting should prioritise risk, not hide it.
Use independent reviewers
Have two reviewers independently score a representative subset. Record agreement, disputed classifications and adjudication. Blind reviewers to the system or version where practical.
Low agreement may indicate an unclear source, weak reference or ambiguous scoring guide and should be reported as a limitation.
Report results by condition
Separate quiet audio, noise, multiple speakers, specialist vocabulary, accents, languages and other relevant conditions. Include file count and audio minutes for each group.
State whether the report uses micro-averaging, which pools all words and errors, or macro-averaging, which gives each file or condition equal weight.
Show variability and uncertainty
Report the median and range across files. Where the sample supports it, include a confidence interval or another transparent uncertainty estimate. Label small studies exploratory rather than implying universal performance.
Set a release rule
Decide whether the workflow passes, passes only with mandatory human checks or fails for the intended use. Link the decision directly to the pre-agreed thresholds and condition-level evidence.
Retain the test-set version, reference rules, settings, raw scoring sheet, reviewer decisions and report. Retest after material changes to hardware, application, processing, language settings or workflow.
Workflow choice matrix for How to Benchmark AI Transcription Accuracy
Choose the method that protects the source and reduces downstream correction. The table makes the non-hardware options explicit.
| Condition | Preferred route | Why |
|---|---|---|
| Repeatable remote work with approved integrations | Cloud software | Automation and central collaboration may outweigh device independence. |
| In-person, mobile or unreliable-connectivity work | Dedicated recorder | Independent capture and a recoverable local source are usually more resilient. |
| Recording is refused, prohibited or unnecessary | Manual notes / no recording | Respecting the boundary is the correct workflow, not a product failure. |
| High-risk or mixed work | Governed hybrid | Separate capture, review, approval and retention rather than trusting one tool. |
Frequently asked questions
Is word error rate enough?
No. It does not reveal whether the system changed a name, amount, negation, decision, deadline or speaker.
Should failed recordings count?
Yes. Excluding failures makes the workflow appear more reliable than it is.
How many recordings are needed?
Enough to represent the intended environments, speakers, languages, durations and vocabulary. Always report the sample size and audio minutes.
Can two vendors use different test sets?
That prevents a fair comparison. Use identical source audio where possible or carefully matched conditions when hardware capture differs.
Useful resources
- NIST AI Risk Management Framework
- How to Compare Two AI Transcription Systems Fairly
- Apply this guide before assessing NERALVO Halo
Benchmark checklist
- Use case and thresholds defined
- Test set locked and versioned
- Reference checked independently
- Failed files included
- Every denominator reconciled
- Critical details, speakers and timestamps measured
- Unsupported additions and severity reported
- Condition-level results and uncertainty shown
- Release rule and retest triggers documented
Related AI voice recorder guides

On this page
Related guides
See whether Halo fits this workflow
Review the NERALVO Halo specifications, included services, delivery information and current offer only after completing the guide.
Found an error or an out-of-date claim? Email support@neralvo.com with the article address and a supporting source.