FREE today: upgraded Halo Productivity Pack worth £29 included with every Halo order
Up to 35 hours recording - 152 languages - 64GB storage
NERALVO
NERALVO
AI recorder guide

How to Benchmark AI Transcription Accuracy by Error Type

Transcript error rates can look precise while still being misleading. The result changes when the reference rules, denominator, sample mix or treatment of failed files changes. A responsible benchmark fixes these choices before testing and reports raw counts, conditions and limitations beside every headline rate.

Define the decision the metric will support

State the intended use of the transcript and what quality is required. A draft internal note may tolerate harmless punctuation differences, while a formal record may require exact names, amounts, dates and speaker attribution.

Set pass, conditional-pass and fail thresholds before viewing results. Combine an overall error measure with specific limits for critical details and processing failures.

Lock and version the test set

Create a manifest of every recording selected for testing. Include sample ID, duration, language, speaker count, room condition, recording mode, device, microphone distance and specialist vocabulary.

Write inclusion and exclusion rules in advance. Preserve corrupted files, failed uploads, empty transcripts and other processing failures in the reconciliation record. Do not quietly remove them from the denominator.

Create a controlled reference transcript

Use a documented transcription convention. Decide how to handle filler words, false starts, punctuation, contractions, partial words, non-speech sounds and inaudible passages.

Have a second reviewer check material words, names, numbers, negations and speaker changes. Lock the reference version before scoring the candidate. If the reference later changes, recalculate the affected results and issue a new report version.

Calculate word error rate correctly

Word error rate is commonly calculated as:

WER = (substitutions + deletions + insertions) ÷ reference words.

For example, a transcript with 22 substitutions, 8 deletions and 5 insertions across 2,500 reference words has a WER of 1.4%.

Report all four raw values: substitutions, deletions, insertions and reference words. WER can exceed 100% when the candidate contains many insertions. It should not automatically be converted into “accuracy” by calculating 100 minus WER without explaining the limitation.

Measure processing-failure rate

A transcription system that fails to return an output should not disappear from the benchmark.

Processing-failure rate = failed files ÷ attempted files.

Define failure before testing. It may include an upload error, unsupported format, empty transcript, incomplete output, unrecoverable timeout or application crash. Report recoverable reruns separately so reliability is visible.

Reconcile every denominator

Keep a table showing:

  • recordings selected;
  • recordings attempted;
  • recordings successfully processed;
  • recordings scored;
  • files or passages excluded;
  • reference words;
  • critical fields;
  • processing failures and reruns.

If the denominator changes between reports, explain why and show the effect on the result where practical.

Measure critical-field accuracy

Create a locked reference list of names, organisations, dates, times, amounts, percentages, units, technical terms, identifiers, negations, decisions, action owners, deadlines and conditions.

Critical-field accuracy = correct critical fields ÷ total reference critical fields.

Also report incorrect, omitted and unsupported fields as raw counts. A field is not correct when the words are similar but the operational meaning has changed.

Measure speaker attribution

Choose one scoring unit: speaking turn, sentence, timed segment or material statement.

Speaker-attribution accuracy = correctly attributed units ÷ total reference units.

Report misattributed and unattributed units separately. A meeting transcript may have a low WER while assigning a commitment to the wrong person.

Measure timestamp accuracy

Define an acceptable tolerance, such as a fixed number of seconds, before testing. Report the percentage of reference points that fall within that tolerance and the median absolute timing error.

Do not judge timestamps by eye after seeing the output.

Measure unsupported additions

For summaries or structured notes, count material claims that have no support in the source.

Unsupported-addition rate = unsupported material assertions ÷ total material assertions in the generated output.

State the assertion unit and reviewer rules. Separate invented decisions, owners or reasons from ordinary word insertions.

Apply a fixed severity framework

Classify errors before scoring:

  • Cosmetic: harmless punctuation, filler or formatting.
  • Operational: requires correction or weakens search and follow-up.
  • High-risk: changes identity, amount, deadline, safety, obligation, decision status or another material outcome.

Publish unweighted counts alongside any severity-weighted score. Weighting can help prioritise risk, but it should not hide the underlying errors.

Use independent reviewers

Have two reviewers independently score a representative subset. Record their agreement, disputed classifications and the adjudication method. Blind reviewers to the system or version where practical.

Low agreement may show that the scoring guide, reference transcript or audio is uncertain. Report it as a limitation.

Report results by condition

Separate quiet audio, background noise, multiple speakers, specialist vocabulary, accents, languages and other relevant operating conditions. Include file count and audio minutes for every group.

Distinguish micro-averaging, which pools all words and errors, from macro-averaging, which gives each file or condition equal weight. State which method is used.

Show variability and uncertainty

Report the median and range across files. Where the sample supports it, include a confidence interval or another transparent uncertainty estimate. A small exploratory test should not be reported with false precision.

Set a release rule

Decide whether the workflow passes, passes only with mandatory human checks or fails for the intended use. Link the decision to the pre-agreed thresholds and condition-level results.

Retain the test-set version, reference rules, source metadata, candidate settings, scoring sheet, reviewer decisions and report. Rerun after material changes to the device, application, model, language settings or workflow.

The NERALVO Halo AI Voice Recorder can provide source audio and DOWAY transcription for a controlled benchmark. Any published rate should name the tested setup, version, sample and scoring method.

Final checklist

  • Are the use case and thresholds defined?
  • Is the test set locked and versioned?
  • Are failed files included transparently?
  • Are raw counts shown beside every rate?
  • Are critical fields, speakers, timestamps and false additions measured?
  • Are reviewer agreement and severity visible?
  • Are condition-level results and uncertainty reported?
  • Can every denominator be reconciled?

Ready to capture meetings properly?

View the NERALVO Halo AI voice recorder with 64GB local storage, meeting capture, compatible phone-call recording workflows and one year of DOWAY Max included.

View NERALVO Halo

Continue reading

Newer guide How to Create a Recording Acceptance Test for Important Workflows Older guide AI Voice Recorder for Compliance Officers: Monitoring Interviews, Findings and Remediation
Browse all AI Recorder Guides articles

Official sources and further reading

Product specifications, policies and legal guidance can change. Check the current official source before making a purchasing, workplace, privacy or compliance decision.