Manuscript Studio

Quality checking

Transcribing the rendered audio and comparing it against your text — catching stage directions read aloud and lines that were never written.

Quality checking is the proofing stage of real audiobook production, and it exists to catch two failure modes that are otherwise completely silent.

The two failures

The engine reads its stage directions aloud. You tagged a line as whispered and the narrator says the word "whispers". This is a documented behaviour of expressive engines, and nothing about the audio file looks wrong.

The engine invents lines. Text that was never in your manuscript, delivered in the same voice as everything around it.

Neither is audible as an error unless you happen to be listening to that exact passage. On a ten-hour book, you will not be.

How it works

Rendered audio is transcribed, and the transcript is compared against the source text.

Two comparisons run, because one is not enough:

  • Similarity catches drift and missing content. It is deliberately tolerant of ordinary transcription variance — "harbor" against "harbour" and the like — because a check that cries wolf gets switched off.
  • A dedicated cue check catches spoken stage directions. This one matters because a passage where the narrator read a cue aloud still scores very high on similarity — almost all the words are correct. Only a check that looks specifically for words you did not write will catch it.

Three results, not two

Result Meaning
Passed Transcript matches the source
Flagged Something to look at, with the passage beside it
Unverified The check could not run

Unverified never displays as a pass. If the transcription itself fails, you are told the passage is unchecked — a quality check that reported success when the checker was broken would be worse than no quality check at all.

Similarly, a passage that has been rendered but never checked shows as rendered, not verified. Those are two different marks and the studio keeps them apart.

Reviewing

Flagged passages are a filter on the script, not a separate screen.

That is deliberate: a flagged passage is only actionable next to its text, where you can see what was supposed to be said, decide whether it matters, and re-render just that passage.

What to watch for

  • Run QC after rendering, before delivery. It is the only step that inspects what the engine actually produced rather than what it was asked to produce.
  • A flag is not automatically a fault. Transcription is imperfect; look at the passage and judge. The check surfaces candidates, it does not adjudicate.
  • Re-rendering a flagged passage is cheap — it is one passage, priced as one passage.