Reviewed and differentiated: 22 July 2026.
Comparing two AI transcription systems is not the same as calculating one system’s abstract accuracy score. A useful A/B comparison asks a practical question: which system produces the safer, faster and more usable result for this organisation’s real recordings?
The test must give both candidates the same source audio, hide their identities from reviewers where possible and include the work required after transcription. A system should not win because it was tested on easier recordings, received more manual help or produced a polished summary that nobody checked.
Start with the buying decision
Define the decision before selecting samples. Examples include:
- Which system should be approved for internal project meetings?
- Which service creates the most usable interview transcript?
- Which workflow reduces correction time without increasing critical errors?
- Which product works reliably across the organisation’s phones, languages and meeting types?
State any non-negotiable conditions, such as successful offline capture, exportable audio, participant controls, UK or EU processing requirements, team administration or a maximum annual cost. A candidate that fails a mandatory condition should not win because its average transcript looks slightly cleaner.
Lock the two candidates and configurations
Record exactly what is being compared:
- product and service name;
- device, phone and operating-system version;
- recording mode and microphone placement;
- application and transcription version;
- language and speaker settings;
- custom vocabulary or glossary;
- subscription tier;
- summary template or prompt;
- export format.
Do not change one candidate’s settings halfway through the test without creating a new labelled test version. Otherwise the final result combines several configurations that cannot be reproduced.
Use one source set for both systems
Both candidates should process the same authorised recordings wherever technically possible. Build a small but representative set containing:
- a quiet one-to-one conversation;
- a normal multi-speaker meeting;
- background noise or room echo;
- overlapping speech;
- names, dates, amounts and reference numbers;
- specialist vocabulary;
- different accents or languages used in real work;
- a longer recording that tests stability and file limits.
Assign each source a neutral sample ID. Preserve the same original audio for both systems so microphone quality or room position does not unfairly determine the winner. Where the products require their own hardware capture, run matched sessions under the same seating, distance, agenda and conditions and disclose that the sources are not identical.
Keep the raw outputs untouched
Save each first transcript and generated summary before anybody corrects it. Record processing time, failures, retries and warnings. Do not silently rerun the weaker candidate until it produces a better result.
Use neutral labels such as System A and System B. Reviewers who know the brand, price or preferred supplier may unconsciously score familiar output more generously.
Compare the complete workflow, not only the transcript
| Dimension | What to measure | Why it matters |
|---|---|---|
| Capture reliability | Files started, completed and recoverable | No transcript exists when capture fails |
| First usable draft | Whether the output is readable and correctly structured | Shows practical starting quality |
| Critical details | Names, figures, dates, negations, owners and deadlines | Small errors can change the outcome |
| Correction effort | Active review and editing time | Determines real productivity value |
| Speaker handling | Correct, wrong and missing attribution | Actions and quotations depend on identity |
| Summary faithfulness | Omissions, invented claims and changed status | Fluent text can still mislead |
| Exports and routing | Audio, transcript, timestamps and destination systems | The result must enter the real workflow |
| Administration | Accounts, permissions, deletion and support | Individual convenience may not scale |
| Total cost | Hardware, plans, labour and rework | Headline subscription price is incomplete |
Use a checked reference for critical passages
Create a human-checked reference for the sections that matter most. It does not need to reproduce every filler word when the intended use is structured meeting notes, but it should accurately preserve all critical statements, decisions, actions and disputed passages.
Mark passages that remain genuinely inaudible or ambiguous. A candidate should not be penalised for refusing to invent words that no reviewer can recover reliably.
Score critical errors separately
Do not let hundreds of correctly transcribed ordinary words hide one changed amount or deadline. For each sample, list the critical fields in advance and record whether each candidate:
- transcribed it correctly;
- transcribed it incorrectly;
- omitted it;
- added something unsupported;
- assigned it to the wrong speaker.
A candidate can pass general readability and still fail the intended workflow because it repeatedly changes material details.
Time the correction process fairly
Give reviewers the same instructions and stopping rule. The timer should include the work needed to reach the agreed quality level, not stop when the text merely looks polished.
Record:
- time opening and organising the file;
- active listening and correction;
- speaker-label repair;
- verification of names and figures;
- summary and action correction;
- export and filing;
- later rework after quality review.
Rotate the order of System A and System B to reduce fatigue and learning effects. A reviewer who hears the same recording twice may remember the content and complete the second version faster.
Judge summaries as separate products
A better verbatim transcript does not automatically create a better summary. Compare summaries against a required-information list containing decisions, actions, owners, deadlines, conditions, risks, dissent and unresolved questions.
Record three failure types:
- Omission: required information is missing.
- Unsupported addition: the summary introduces a claim not supported by the source.
- Changed meaning: a proposal becomes a decision, a condition disappears or an owner changes.
For a detailed method, use the separate guide on auditing AI meeting summaries for missing information.
Test reliability and recovery
Include at least one controlled failure scenario for each candidate:
- interrupted transfer;
- poor connectivity;
- long recording;
- low device storage;
- application restart;
- wrong account or expired allowance;
- failed processing or empty transcript.
Record whether the original audio remains safe, whether the user receives a clear warning and how much work recovery requires. A system that produces excellent transcripts only when everything works may be less suitable than a slightly less accurate system with reliable recovery.
Compare privacy and operational control
Create a side-by-side data-path map for both systems:
- Where is audio created?
- When is it uploaded?
- Who processes it?
- Who can access audio and text?
- Which copies are created through exports or integrations?
- How are users removed?
- How is deletion performed and evidenced?
- What remains after cancellation?
Do not award a generic “secure” score based only on marketing claims. Judge whether each workflow satisfies the organisation’s actual requirements.
Calculate the comparable cost
Use the same expected monthly volume for both systems. Include:
- hardware and replacement cost;
- subscription or processing allowance;
- required accessories;
- active correction and review labour;
- administration and training;
- failed-file recovery;
- storage, export or integration costs;
- renewal price after any included first year.
A cheaper plan can become the more expensive workflow when staff spend substantially longer correcting and routing its output.
Use weighted criteria only after setting hard gates
First apply mandatory pass/fail gates, such as:
- no unrecoverable source-audio loss;
- zero tolerance for specified high-risk errors;
- required export format available;
- approved participant and data controls;
- compatibility with the organisation’s devices;
- cost within the approved ceiling.
Then weight the remaining criteria according to the use case. A research team may value quotation accuracy and source timestamps more heavily, while a sales team may prioritise action routing and CRM integration.
| Example criterion | Possible weight |
|---|---|
| Critical-detail accuracy | 30% |
| Correction time | 20% |
| Capture and processing reliability | 15% |
| Summary usefulness | 15% |
| Exports and integration | 10% |
| Administration and support | 5% |
| Total cost | 5% |
These are examples, not universal weights. Publish the selected weights before calculating the winner.
Choose by use case when there is no universal winner
System A may win quiet online meetings while System B performs better in noisy in-person rooms. Do not force one overall winner when the evidence supports separate approvals.
A defensible result may be:
- approve A for online internal meetings;
- approve B for field interviews;
- require human verification for both;
- reject both for specified high-risk records;
- run a longer pilot because the sample is too small.
Use a standard comparison report
The final report should include:
- decision and intended use;
- candidate versions and configurations;
- source-set manifest;
- mandatory gates;
- raw output and failure log;
- critical-detail results;
- correction-time results;
- summary findings;
- privacy, administration and cost comparison;
- condition-level strengths and weaknesses;
- reviewer disagreements;
- approved scope, restrictions and retest triggers.
Where NERALVO Halo fits
NERALVO Halo can be one candidate in a controlled A/B comparison. It provides 64GB local storage, NOTE mode, supported CALL recording subject to compatibility and up to 35 hours of recording under suitable conditions. DOWAY provides transcription and structured outputs.
Do not give Halo or any competitor favourable treatment because of brand familiarity. Test the exact purchased configuration against the same real workload and disclose NERALVO’s commercial interest.
Final checklist
- Is the buying decision defined?
- Are both configurations locked and documented?
- Did both candidates receive the same or properly matched source material?
- Were raw outputs preserved before correction?
- Were reviewers blinded where practical?
- Were critical details and summaries checked separately?
- Was total correction effort measured?
- Were failure recovery, privacy, administration and cost included?
- Were mandatory gates applied before weighted scoring?
- Does the final approval reflect condition-specific evidence rather than one flattering average?
Ready to capture meetings properly?
View the NERALVO Halo AI voice recorder with 64GB local storage, meeting capture, compatible phone-call recording workflows and one year of DOWAY Max included.
View NERALVO Halo