Case StudyPrivate Healthcare Clinic

MediScan: AI Medical Assistant

A real-time voice agent that conducts structured medical interviews, analyzes scans, triages urgency, and escalates critical cases to a physician via automated email handoff.

Voice AI
RAG
Healthcare
MediScan: AI Medical Assistant

Project Overview

MediScan-V2 is an AI-powered medical assistant that bridges the gap between patients seeking preliminary guidance and doctors needing structured, prioritized consultation reports. A real-time voice agent conducts a full medical interview, analyzes uploaded scans and lab reports, determines triage urgency, and for critical cases emails a structured clinical summary straight to a supervising physician.

The Problem

Telemedicine intake is slow and inconsistent. Patients struggle to describe symptoms clearly, doctors receive unstructured notes, and genuinely urgent cases can sit in the same queue as routine questions until a human finally reads them.

The Goal

Build an empathetic voice agent that conducts a structured interview, grounds its reasoning in real medical evidence, triages urgency reliably, and gets critical cases in front of a doctor immediately without the patient having to type a word.

What Was Built

1. Real-Time Voice Assistant

LiveKit and Google Gemini Live power sub-second, full-duplex voice conversation, with Silero VAD for accurate speech detection and natural conversational turn-taking.

2. Intelligent Medical Triage

The agent autonomously tracks conversation stage (Initial, Gathering, Clarifying, Concluding) to ensure a complete clinical picture, with built-in guardrails that deflect off-topic queries and keep the interview focused on healthcare.

3. Evidence-Based RAG

A Pinecone vector database, queried with the localized MedEmbed embedding model, maps patient symptoms against verified medical cases and guidelines to ground every response in literature.

4. Medical Scan Analysis

Patients upload brain, kidney, chest, or breast scans. The AI extracts key findings, confidence scores, and clinical significance, then attaches its recommendation for professional review.

5. Automated Doctor Handoff

When triage reaches EMERGENCY or URGENT, or the patient asks, the agent synthesizes the conversation into a structured clinical report and emails it to a supervising physician via SMTP.

System Architecture

MediScan-V2 runs as a decoupled microservices architecture: a Next.js frontend, a FastAPI backend, and a persistent LiveKit Agent Worker that owns the live voice conversation from greeting to handoff.

The flow moves through six stages:

  • User Login to Room Connect: User authenticates and a LiveKit room is created
  • Agent Dispatch: FastAPI generates a secure connection token and dispatches the job
  • Voice Consult: The agent conducts a full medical interview using Gemini Live
  • Knowledge Retrieval: Complex symptoms trigger a vector query to ground the next question
  • Triage Assess: Every turn silently re-evaluates urgency level and conversation stage
  • Doctor Handoff: Critical triage compiles a clinical summary and emails it via SMTP

Scan Analysis

Beyond voice, patients can jump straight to scan analysis. The same evidence-grounded reasoning that powers the voice agent's RAG pipeline backs every image-based diagnosis.

  • Brain Scan: Detects tumors (Pituitary, Meningioma, Glioma) from MRI images
  • Kidney Scan: Detects tumors in the kidney from CT scan images
  • Chest X-Ray: Screens for lung conditions and pneumonia
  • Breast Scan: Detects breast cancer from histopathology microscopy images

Every scan result follows the same structure: a clear triage label with confidence score, a findings checklist in plain clinical language, and a concrete next step always paired with a professional review disclaimer.

Problem Solving in Practice

Keeping a Voice Agent Both Fast and Clinically Safe

A consultation has to feel like a natural conversation, not a form, but every turn also needs to silently re-evaluate conversation stage, triage level, and whether the topic has drifted off-medical. The fix was running Gemini Live over LiveKit for sub-second full-duplex audio while a separate structured-output LangChain task tracks stage and triage in the background, so the safety layer never adds a perceptible pause.

Grounding Symptom Reasoning in Real Medical Evidence

An LLM that free-associates about symptoms is a liability in healthcare. Every non-trivial symptom mention needed to be checked against real medical literature before the agent's next question or recommendation. The fix was a Pinecone-backed RAG layer using domain-specific MedEmbed embeddings, so retrieval reflects clinical similarity rather than generic semantic similarity.

Key Learnings

  • Latency is a trust signal in healthcare UX. Sub-second, full-duplex voice turned out to matter as much for perceived trustworthiness as for usability
  • Guardrails belong in the conversation loop, not bolted on after. Off-topic detection and stage tracking run on every turn, in the background
  • Generic embeddings are not good enough for clinical retrieval. Swapping a general-purpose transformer for MedEmbed was the difference between reassuringly relevant and web-search-level results
  • Escalation has to be automatic, not optional. Triggering doctor handoff at EMERGENCY/URGENT thresholds is what actually protects patients who do not know how serious their symptoms are

Tech Stack

  • Voice Orchestration: LiveKit Agents framework with a persistent Python agent worker
  • Real-Time Audio: Gemini Live native-audio model with Silero VAD
  • Retrieval: Pinecone vector search with domain-specific MedEmbed embeddings
  • State Management: Concurrent ConversationState tracking demographics, language, and triage
  • Auth and Frontend: Next.js 16, React 19, TailwindCSS v4, Supabase, LiveKit React components
  • Doctor Handoff: FastAPI with SMTP for structured clinical report generation and delivery
building something exciting?

We want to hear from you

Join 50+ businesses scaling with AI.

Powered By

Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make

Book a technical discovery

Tell us about your operational bottlenecks. We'll engineer an autonomous system to eliminate them entirely.