Building Voice Auditor: the polite call that broke two rules
Sep 28, 2026 · 10 min read

The call is 45 seconds long. A customer has been charged twice, and the conversation stays polite from the first line to the last. Twenty-four seconds in, the agent says this:
It won't. And if you move to our SAVER plan, it's guaranteed to lower your bill every month.
Nothing about it sounds wrong. The models I run on every call agreed: positive sentiment, gratitude as the strongest emotion, toxicity close to zero, and a compliance score of 99.9 out of 100. But the company's policy says an agent must never promise a guaranteed result, and must tell the caller that the call may be recorded. This call did neither. It broke two rules and sounded perfect doing it.
That gap, between how a call sounds and what was actually said, is what Voice Auditor is built around.
Why I built it
Most of a compliance review is listening. Someone plays each recording at normal speed with the policy open beside it and checks the same things every time. Was the recording disclosed? Did anyone promise what they shouldn't? Did the conversation stay civil? It is slow, it is tiring, and two reviewers can hear the same call differently. The rules themselves live in a document, and a document can't flag a single call.
I wanted to see how much of that first pass open models could take on, and to turn the policy into rules that a system checks on every call, the same way each time. The call above is one of ten I scripted for a made-up energy company and voiced with macOS text-to-speech. Synthetic calls let me publish every second of the run, and the film on the project page is a real recording of the app working through them.
What happens to one call
| Step | What happens |
|---|---|
| Upload | A file, or a recording made in the browser. |
| Transcribe | Whisper, the small model, writes a timestamped transcript and detects the language. |
| Redact | Card numbers, emails, phone numbers and similar details become placeholders. Nothing after this step sees them. |
| Analyze | Open models score sentiment, emotion and toxicity, classify intent, find topics and write a summary. |
| Score | Sentiment, emotion and toxicity combine into a score from 0 to 100. |
| Check | The team's own rules run against the transcript and the scores. |
| Notify | Findings go out by webhook and email, and the call is saved for review. |
The backend is FastAPI, and one request runs the whole chain. Each step is wrapped on its own, so a model that fails leaves a gap in the report instead of failing the call. Every model is open. Sentiment comes from a RoBERTa that Cardiff NLP trained on tweets, and emotion from a RoBERTa fine-tuned on Google's GoEmotions. Detoxify scores toxicity. BART writes the summary and classifies intent zero-shot, LDA finds the topics, and SHAP shows which words pushed the toxicity score up. The front end is React and TypeScript, with history, statistics, side-by-side comparison, comments, tags and roles, so a team can review the same calls together.
None of it needs a GPU. On my laptop's CPU an analysis takes about twice as long as the call itself: the ten test calls averaged 30 seconds of audio and 61 seconds of analysis.
Scores hear tone. Rules read words.
The score is simple on purpose:
score = 100 × sentiment weight × emotion penalty × (1 − toxicity)
sentiment weight negative 0.6, neutral 0.9, positive 1.0
emotion penalty anger 0.6, fear 0.8, sadness 0.8, anything else 1.0
A call with an abusive agent and an angry customer will score low, which is useful. But the score measures tone, and policy is mostly about words. The guarantee in the call above was delivered warmly, so the score had nothing to catch.
Catching it is the job of the rules. A team writes each rule once, gives it a severity, and can test it against sample text. There are seven kinds: regex, keyword, sentiment, emotion, toxicity, score threshold and custom. My test run used four:
guarantee language regex \bguarantee(d|s)?\b
risk-free promises keyword risk-free, promise you
recording disclosed custom "recorded" not in text.lower()
stay civil toxicity > 0.5
Across the ten calls, the rules flagged three. The call above broke two rules at once. A refund call made a risk-free promise, and in a late-fee dispute nobody mentioned the recording. The score flagged none. Nine calls scored exactly 99.9 and the tenth 87.8, and part of the reason is a bug I'll come to.
A finding doesn't wait for someone to open the app. The guarantee rule is marked critical, the missing disclosure a warning, and both were posted to a QA alerts webhook as the analysis finished.
What the first version got wrong
The first version had a longer feature list than its code could honestly support. When I went back through it with a reviewer's eye, four problems stood out.
Sentiment never counted. The sentiment model names its classes LABEL_0, LABEL_1 and LABEL_2. My weight table looked up NEGATIVE, NEUTRAL and POSITIVE, found nothing, and fell back to 1.0 on every call. That same mismatch meant negative-sentiment alerts never fired and sentiment rules never matched. The front end hid all of it by translating the labels for display. That is a large part of why nine calls scored 99.9. The fix renames the model's classes once, when it loads, so everything downstream gets words, and a migration rewrites the old labels in stored records. Scores now move the way the weights intended: a neutral test call that used to score 99.9 now scores 89.9.
Custom rules ran through eval. The first version evaluated them like this:
eval(rule.pattern, {"__builtins__": {}}, context)
Empty builtins look safe, and they aren't. Python lets you walk from any object to its class and on to every class loaded in the process, so a rule as short as ().__class__.__bases__[0].__subclasses__() reaches far outside its sandbox. Rules now go through a small interpreter that walks the parsed expression and allows only what a rule needs. That means comparisons, and/or/not, basic arithmetic, the rule's own variables, a few functions such as contains and matches, and read-only text and dictionary methods. It refuses **, repeating a string a billion times, and any name that starts with an underscore. A rule is checked when it is saved, so its author finds out straight away, and again every time it runs.
Most of the API was open. 38 of the 64 routes needed no sign-in. The first administrator was admin with the password admin123, and tokens were signed with a fallback secret that sat in the public repository. Now every route needs a signed-in user with a named permission, and there is no default password: the first sign-in forces a new one. The published example secrets are refused, sign-ins are rate limited, and every sign-in and change is written to an audit log with the account, the outcome and the address it came from.
Speaker turns are a placeholder. The audio is cut into two-second windows, and each window is labelled by its spectral centroid. That is not diarization, and the project page says so. It is the next thing I will replace.
Keeping card numbers out of the database
Support calls are full of things that should never be stored: card numbers read out digit by digit, security codes, dates of birth, email addresses. Voice Auditor now replaces them straight after transcription, before any model, rule, database row, webhook or email sees the text.
The detector is pattern matching with validation:
- Card numbers must pass the Luhn check, and IBANs the mod-97 check.
- Social Security numbers must fall outside the ranges that were never issued.
- Security codes, PINs, account numbers and dates of birth count only when the words just before them say what they are.
- Digits spoken as words are matched too.
It is fast enough that nobody notices: a 39,000-character transcript takes about 50 milliseconds.
My first version passed every test I wrote for it. Then I generated a billing call in which the customer reads out a test card number, an expiry date, a security code, an email address and a phone number, and ran it through Whisper. Across two runs, this is what came back:
| What was said | What Whisper wrote |
|---|---|
4111 1111 1111 1111 | 4111,111,111,111 |
09/28 | 09 strobe 28, and on the next run, 0 9Strogue 28 |
jane.doe@example.com | jane.doe. At example, dot com, and then Jane.doe.at. Example.com |
The card number came back grouped like an amount and three digits short. It failed the Luhn check, so it walked straight through. The email and the expiry got through too. None of those strings were in my tests, because I had written the tests from what I expected a transcript to look like.
So the tests now start from what Whisper actually writes. After words such as "card", "credit" or "Visa", any run of 12 to 19 digits is treated as a card number, whatever the checksum says. A card number with one digit wrong is still a card number. The email and expiry patterns tolerate the punctuation and mishearings Whisper adds, and each of those outputs is now a test case. On the third run, all five were replaced:
It is [CARD]. The expiration date is [EXPIRY] and the security code is [CVV].
... My email is [EMAIL] and my phone number is [PHONE]
A scan of the whole response and of the stored record found none of the original values. Each analysis keeps a count of what was removed, and the count raises an alert, because a customer reading a card number aloud on a recorded line is something a compliance team needs to know about. In the app, the placeholders show up as small lock chips in the transcript.
It is still pattern matching, and speech-to-text will keep finding new ways to write a number, so I treat it as a strong safeguard rather than a guarantee. Recordings get the same care. Each upload is deleted as soon as its analysis ends, where the first version deleted it only when the analysis succeeded, so every failure left a recording on disk. Analyses and audit entries can also be set to expire after a number of days.
Smaller lessons
The API used to take minutes to start. One module loaded two models the moment it was imported, and three modules each kept their own copy of the same models. Every model now loads on first use, once per process, and the API starts in about a second.
While making that change, I introduced a bug of my own. To load a model lazily, I handed SHAP a stand-in object instead of the real tokenizer. SHAP decides how to explain a model by looking at the tokenizer's class, so it treated the transcript as numbers and crashed on the first real call. All 48 tests passed, because none of them loaded a model. There is now a set of model smoke tests that loads and runs every model an analysis uses, and I run them after any model or dependency change.
I also added dependency audits to the build. The first checks turned up known vulnerabilities in several pinned Python packages, and 27 advisories in the front end's lock file, which hadn't changed since the first commit. One was an open-redirect issue in the version of React Router the app shipped. After the upgrades both audits are clean, and a new high-severity advisory now fails the build instead of scrolling past in a log.
What's next
Three things are next:
- Real speaker diarization, so "who said it" is as reliable as "what was said".
- A set of labelled calls to calibrate the score. The weights came from the first version, and I doubt a neutral call is really ten percent less compliant than a cheerful one.
- A job queue and separate workspaces, once it runs for a team rather than on my laptop. The queue means a long call no longer holds a request open for a minute, and each team gets its own workspace.
Try it
The project page has an 84-second film of the run above, with a text version. The code is open source under the MIT license on GitHub, with setup steps in the README. If your team still reviews calls by listening to them one at a time, I'd like to hear how it goes: abhijeet.solanki@outlook.com.