An AI agent that telephones a patient, confirms who it is speaking to before saying anything clinical, reads out their recent test results, and books a follow-up appointment through a tool call — with full observability and an automated privacy check scoring every call.
A full recorded demo: two real outbound phone calls, the agent's live view, and the resulting traces in Opik.
VerifiedAgent — followed by the post-call analysis and the Opik trace.Automated healthcare calls carry a privacy problem that most voice agents quietly ignore: the system placed the call, so it has no idea who picked up. A partner, a flatmate or a wrong number can answer a patient's phone. And on an outbound call every question leaks information — even an innocent-sounding "I'm calling about your recent test results" has already disclosed that this person is a patient of the clinic who has had tests done, before a single number is read aloud.
A prompt instruction is not a real defence here. It is a soft constraint that a model can drift past under pressure from an insistent caller, a confused patient, or a deliberate prompt injection.
Build a production-shaped outbound voice agent that holds a natural phone conversation over the real telephone network, discloses health data only to a verified patient, books an appointment through a genuine tool call, and can prove afterwards that it behaved correctly — with an observability layer that plugs in without the agent code knowing it exists.
UnverifiedAgent is constructed with no health data and no booking tools; on a successful check the verification tool returns a VerifiedAgent instance, which triggers a framework handoff while the audio session continues uninterrupted. The agent cannot disclose what it was never constructed with — a guarantee you verify by reading a constructor, not by trusting a model.nova-3-medical, GPT-4.1-mini, Inworld TTS) over a speech-to-speech model, so every turn and every tool call stays inspectable — which is what the observability and evaluation layers are built on.VerifiedAgent is never constructed: no biomarker appears anywhere in the transcript, and there is no booking span in the trace. The exported traces are in the repository so the claim can be checked directly.