MediScan: AI Medical Assistant
A real-time voice agent that conducts structured medical interviews, analyzes scans, triages urgency, and escalates critical cases to a physician via automated email handoff.

Project Overview
MediScan-V2 is an AI-powered medical assistant that bridges the gap between patients seeking preliminary guidance and doctors needing structured, prioritized consultation reports. A real-time voice agent conducts a full medical interview, analyzes uploaded scans and lab reports, determines triage urgency, and for critical cases emails a structured clinical summary straight to a supervising physician.
The Problem
Telemedicine intake is slow and inconsistent. Patients struggle to describe symptoms clearly, doctors receive unstructured notes, and genuinely urgent cases can sit in the same queue as routine questions until a human finally reads them.
The Goal
Build an empathetic voice agent that conducts a structured interview, grounds its reasoning in real medical evidence, triages urgency reliably, and gets critical cases in front of a doctor immediately without the patient having to type a word.
What Was Built
1. Real-Time Voice Assistant
LiveKit and Google Gemini Live power sub-second, full-duplex voice conversation, with Silero VAD for accurate speech detection and natural conversational turn-taking.
2. Intelligent Medical Triage
The agent autonomously tracks conversation stage (Initial, Gathering, Clarifying, Concluding) to ensure a complete clinical picture, with built-in guardrails that deflect off-topic queries and keep the interview focused on healthcare.
3. Evidence-Based RAG
A Pinecone vector database, queried with the localized MedEmbed embedding model, maps patient symptoms against verified medical cases and guidelines to ground every response in literature.
4. Medical Scan Analysis
Patients upload brain, kidney, chest, or breast scans. The AI extracts key findings, confidence scores, and clinical significance, then attaches its recommendation for professional review.
5. Automated Doctor Handoff
When triage reaches EMERGENCY or URGENT, or the patient asks, the agent synthesizes the conversation into a structured clinical report and emails it to a supervising physician via SMTP.
System Architecture
MediScan-V2 runs as a decoupled microservices architecture: a Next.js frontend, a FastAPI backend, and a persistent LiveKit Agent Worker that owns the live voice conversation from greeting to handoff.
The flow moves through six stages:
- User Login to Room Connect: User authenticates and a LiveKit room is created
- Agent Dispatch: FastAPI generates a secure connection token and dispatches the job
- Voice Consult: The agent conducts a full medical interview using Gemini Live
- Knowledge Retrieval: Complex symptoms trigger a vector query to ground the next question
- Triage Assess: Every turn silently re-evaluates urgency level and conversation stage
- Doctor Handoff: Critical triage compiles a clinical summary and emails it via SMTP
Scan Analysis
Beyond voice, patients can jump straight to scan analysis. The same evidence-grounded reasoning that powers the voice agent's RAG pipeline backs every image-based diagnosis.
- Brain Scan: Detects tumors (Pituitary, Meningioma, Glioma) from MRI images
- Kidney Scan: Detects tumors in the kidney from CT scan images
- Chest X-Ray: Screens for lung conditions and pneumonia
- Breast Scan: Detects breast cancer from histopathology microscopy images
Every scan result follows the same structure: a clear triage label with confidence score, a findings checklist in plain clinical language, and a concrete next step always paired with a professional review disclaimer.
Problem Solving in Practice
Keeping a Voice Agent Both Fast and Clinically Safe
A consultation has to feel like a natural conversation, not a form, but every turn also needs to silently re-evaluate conversation stage, triage level, and whether the topic has drifted off-medical. The fix was running Gemini Live over LiveKit for sub-second full-duplex audio while a separate structured-output LangChain task tracks stage and triage in the background, so the safety layer never adds a perceptible pause.
Grounding Symptom Reasoning in Real Medical Evidence
An LLM that free-associates about symptoms is a liability in healthcare. Every non-trivial symptom mention needed to be checked against real medical literature before the agent's next question or recommendation. The fix was a Pinecone-backed RAG layer using domain-specific MedEmbed embeddings, so retrieval reflects clinical similarity rather than generic semantic similarity.
Key Learnings
- Latency is a trust signal in healthcare UX. Sub-second, full-duplex voice turned out to matter as much for perceived trustworthiness as for usability
- Guardrails belong in the conversation loop, not bolted on after. Off-topic detection and stage tracking run on every turn, in the background
- Generic embeddings are not good enough for clinical retrieval. Swapping a general-purpose transformer for MedEmbed was the difference between reassuringly relevant and web-search-level results
- Escalation has to be automatic, not optional. Triggering doctor handoff at EMERGENCY/URGENT thresholds is what actually protects patients who do not know how serious their symptoms are
Tech Stack
- Voice Orchestration: LiveKit Agents framework with a persistent Python agent worker
- Real-Time Audio: Gemini Live native-audio model with Silero VAD
- Retrieval: Pinecone vector search with domain-specific MedEmbed embeddings
- State Management: Concurrent ConversationState tracking demographics, language, and triage
- Auth and Frontend: Next.js 16, React 19, TailwindCSS v4, Supabase, LiveKit React components
- Doctor Handoff: FastAPI with SMTP for structured clinical report generation and delivery