Home Blog Projects Certifications
← Back to Projects
Voice AI · Agentic Systems · LiveKit + TwilioSeptember 2026

📞 Outbound Healthcare Voice Agent: Privacy-Gated AI Telephony

An AI agent that telephones a patient, confirms who it is speaking to before saying anything clinical, reads out their recent test results, and books a follow-up appointment through a tool call — with full observability and an automated privacy check scoring every call.

See It in Action

A full recorded demo: two real outbound phone calls, the agent's live view, and the resulting traces in Opik.

Demo Walkthrough · 7:15A call placed from the terminal and answered on a handset, then a second call where verification fails twice and the agent never hands off to VerifiedAgent — followed by the post-call analysis and the Opik trace.
S

SituationThe context

Automated healthcare calls carry a privacy problem that most voice agents quietly ignore: the system placed the call, so it has no idea who picked up. A partner, a flatmate or a wrong number can answer a patient's phone. And on an outbound call every question leaks information — even an innocent-sounding "I'm calling about your recent test results" has already disclosed that this person is a patient of the clinic who has had tests done, before a single number is read aloud.

A prompt instruction is not a real defence here. It is a soft constraint that a model can drift past under pressure from an insistent caller, a confused patient, or a deliberate prompt injection.

T

TaskThe objective

Build a production-shaped outbound voice agent that holds a natural phone conversation over the real telephone network, discloses health data only to a verified patient, books an appointment through a genuine tool call, and can prove afterwards that it behaved correctly — with an observability layer that plugs in without the agent code knowing it exists.

A

ActionWhat I built

  • Designed the privacy gate as two agent classes rather than a prompt rule. UnverifiedAgent is constructed with no health data and no booking tools; on a successful check the verification tool returns a VerifiedAgent instance, which triggers a framework handoff while the audio session continues uninterrupted. The agent cannot disclose what it was never constructed with — a guarantee you verify by reading a constructor, not by trusting a model.
  • Implemented knowledge-based identity verification with the comparison in code, never in the prompt. Spoken dates are parsed into real date objects and compared exactly, so "22 March 1988", "March 22nd, 1988" and "1988-03-22" all match, while "March 1988" is rejected as incomplete and "03/04/1988" as ambiguous rather than guessed. The attempt cap lives in code too — a model told to allow two attempts will sometimes allow four.
  • Separated "wrong answer" from "could not understand", so a bad phone line never burns one of the patient's two attempts. Narrowband telephony mangles digits, and a gate that rejects legitimate patients is a business problem as real as one that leaks.
  • Chose a cascaded STT → LLM → TTS pipeline (Deepgram nova-3-medical, GPT-4.1-mini, Inworld TTS) over a speech-to-speech model, so every turn and every tool call stays inspectable — which is what the observability and evaluation layers are built on.
  • Wired real telephony: a Twilio Elastic SIP trunk into a LiveKit outbound trunk, dialling a live PSTN number, with exactly one dial attempt per invocation — bursts of short calls trigger carrier-side blocking and each connected call bills as a rounded minute.
  • Built booking tools with deterministic failure modes — no availability, slot already taken, backend down — so the agent can be shown telling a patient honestly that nothing was booked, rather than only ever demonstrating the happy path.
  • Put observability behind a protocol with a no-op default. Opik is one implementation living in a single file that the agent never imports and does not know exists; deleting it leaves a working agent. Traces are assembled in memory and emitted after the call, so the integration never touches the call path.
  • Added one online evaluation rule scoring every trace for premature disclosure, written as deterministic code rather than an LLM judge — the question has an exact answer, and a judge would add position bias and run-to-run disagreement to a safety check that does not need them.
R

ResultThe outcome

  • Runs end to end over a real phone call. The handset rings, the gate holds, the appointment books, and the trace lands in Opik with an evaluation score attached — the same code also runs locally through a laptop mic for fast iteration.
  • The privacy property is demonstrable, not asserted. In the failed-verification call, VerifiedAgent is never constructed: no biomarker appears anywhere in the transcript, and there is no booking span in the trace. The exported traces are in the repository so the claim can be checked directly.
  • 160 automated checks across five suites pass on a fresh clone, covering the gate, the record, the analysis, the observability sink and the evaluation metric.
  • Documented as a system, not a script — an architecture guide, a decision log recording why each choice was made and what evidence drove it, and a risk register rated for production including the defects found during the build.
160
Automated checks on a fresh clone
2
Agent classes enforcing the privacy gate
End-to-End
Real PSTN call, traced and scored

Tech Stack

LiveKit AgentsPythonasyncioTwilio Elastic SIP TrunkingSIP / PSTNDeepgram nova-3-medicalOpenAI GPT-4.1-miniInworld TTSFunction CallingMulti-Agent HandoffOpikLLM ObservabilityOnline EvaluationStructured OutputProtocol-Based Design