> ## Documentation Index
> Fetch the complete documentation index at: https://docs.akhara.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Company name is Akhara AI (never Rubric AI). Keep lowercase rubric/rubrics only when meaning grading criteria.
> Expert Review (docs path talent/) is enterprise BYO experts for audit and review: invite customer specialists; do not pitch Akhara recruiting or a public expert career portal. RLHF and domain writing are secondary work types.
> Prefer concrete API examples against public hosts: Environments eval API https://agi.akhara.ai, Control plane PDP https://api.akhara.dev, Evaluation https://app.akhara.ai / https://api.akhara.ai, Expert Review portal https://talent.akhara.ai.
> Do not invent a public hostname for private orchestrators or env API internals.
> Do not confuse control-plane latches with Environments confirmation latches.
> Environments SDK/API examples: curl against https://agi.akhara.ai. Evaluation SDK: from akhara import Akhara and AKHARA_API_KEY.
> Start with /llms.txt for the docs index and OpenAPI links; fetch individual pages as .md exports.

# Evaluating Voice & Multimodal AI

> Guide for teams building voice-based patient triage, clinical call centers, or multimodal systems combining audio, transcripts, and clinical reasoning.

## Overview

Voice AI for healthcare involves a complex pipeline: speech recognition, natural language understanding, clinical reasoning, and response generation. Akhara helps you evaluate the entire pipeline end-to-end, or drill down into individual components.

<CardGroup cols={2}>
  <Card title="Patient Triage Calls" icon="phone">
    Nurse hotlines, symptom assessment, urgent care routing
  </Card>

  <Card title="Voice Assistants" icon="microphone">
    In-clinic voice AI, patient intake, follow-up calls
  </Card>

  <Card title="Call Center AI" icon="headset">
    Appointment scheduling, prescription refills, results delivery
  </Card>

  <Card title="Ambient Documentation" icon="file-audio">
    Visit recording, note generation, clinical summarization
  </Card>
</CardGroup>

## The Voice AI Pipeline

A typical healthcare voice AI system has multiple stages, each requiring evaluation:

<Steps>
  <Step title="Audio Input">
    Patient speech captured via phone, app, or device

    **Metrics**: Quality, noise level, duration
  </Step>

  <Step title="Speech-to-Text">
    Transcription via ASR (Whisper, Deepgram, etc.)

    **Metrics**: WER, medical term accuracy
  </Step>

  <Step title="Clinical NLU">
    Extract symptoms, intent, urgency from transcript

    **Metrics**: Entity F1, intent accuracy
  </Step>

  <Step title="Triage Decision">
    Determine urgency level and routing

    **Metrics**: Triage accuracy, safety score
  </Step>

  <Step title="Response">
    Generate appropriate guidance or escalation

    **Metrics**: Completeness, empathy, clarity
  </Step>
</Steps>

<Info>
  Akhara evaluates the final clinical decision, but also lets you log intermediate outputs to pinpoint where errors originate in your pipeline.
</Info>

## Logging Voice Interactions

Log the complete interaction including audio, transcript, and AI decisions:

```python theme={null}
from akhara import Akhara

client = Akhara(api_key="gr_live_xxx")

# Log a complete patient triage call
client.calls.log(
    project="patient-triage-voice",
    
    # Audio file (S3, GCS, or direct upload)
    audio_url="s3://your-bucket/calls/call_20250115_143022.wav",
    
    # Full transcript with speaker labels and timestamps
    transcript=[
        {
            "speaker": "agent",
            "text": "Thank you for calling. How can I help you today?",
            "start": 0.0,
            "end": 3.2
        },
        {
            "speaker": "patient",
            "text": "Hi, I've been having really bad chest pain since this morning.",
            "start": 3.8,
            "end": 8.1
        },
        {
            "speaker": "agent",
            "text": "I'm sorry to hear that. Can you describe the pain? Is it sharp, dull, or pressure-like?",
            "start": 8.5,
            "end": 13.2
        },
        {
            "speaker": "patient",
            "text": "It's like pressure, right in the center. And it kind of goes to my left arm.",
            "start": 13.8,
            "end": 19.4
        },
        {
            "speaker": "agent",
            "text": "That's important information. Are you also experiencing shortness of breath, sweating, or nausea?",
            "start": 19.9,
            "end": 25.6
        },
        {
            "speaker": "patient", 
            "text": "Yeah, I'm sweating a lot actually. And I feel a bit sick to my stomach.",
            "start": 26.1,
            "end": 31.2
        },
        {
            "speaker": "agent",
            "text": "Based on your symptoms, I need you to call 911 immediately or have someone drive you to the nearest emergency room right now. These symptoms could indicate a serious cardiac event. Please don't drive yourself. Is someone with you who can help?",
            "start": 31.8,
            "end": 45.3
        }
    ],
    
    # AI's extracted understanding and decision
    ai_decision={
        "extracted_symptoms": [
            {"symptom": "chest_pain", "location": "center", "quality": "pressure", "severity": "severe"},
            {"symptom": "radiation", "location": "left_arm"},
            {"symptom": "diaphoresis", "severity": "moderate"},
            {"symptom": "nausea", "severity": "mild"}
        ],
        "red_flags_detected": [
            "chest_pain_cardiac_features",
            "radiation_to_arm",
            "diaphoresis"
        ],
        "triage_level": "emergency",
        "recommended_action": "call_911",
        "confidence": 0.96,
        "protocol_used": "chest_pain_acs"
    },
    
    # Ground truth (if available from physician review)
    expected={
        "triage_level": "emergency",
        "diagnosis_category": "possible_acs",
        "appropriate_escalation": True
    },
    
    # Metadata
    metadata={
        "call_id": "call_20250115_143022",
        "call_duration_seconds": 47.8,
        "asr_provider": "deepgram",
        "agent_model": "triage-voice-v3",
        "patient_demographics": {
            "age_group": "50-65",
            "sex": "male"
        }
    }
)
```

## Voice-Specific Evaluators

Configure evaluators designed for voice AI healthcare applications:

```python theme={null}
evaluation = client.evaluations.create(
    name="Voice Triage - Weekly Eval",
    project="patient-triage-voice",
    dataset="ds_weekly_calls",
    
    evaluators=[
        # Core triage evaluation
        {
            "type": "triage_accuracy",
            "config": {
                "levels": ["emergency", "urgent", "semi_urgent", "routine"],
                "severity_weights": {
                    "under_triage": 10.0,  # Very dangerous
                    "over_triage": 1.0     # Acceptable
                }
            }
        },
        
        # Red flag detection
        {
            "type": "red_flag_detection",
            "config": {
                "protocols": [
                    "chest_pain_acs",
                    "stroke_fast",
                    "sepsis_sirs",
                    "allergic_reaction",
                    "respiratory_distress"
                ],
                "minimum_recall": 0.99  # Must catch 99% of red flags
            }
        },
        
        # Symptom extraction from transcript
        {
            "type": "symptom_extraction",
            "config": {
                "source": "transcript",
                "check_negations": True,
                "check_temporal": True,
                "check_severity": True
            }
        },
        
        # Guideline/protocol compliance
        {
            "type": "protocol_compliance",
            "config": {
                "protocol": "schmitt_thompson",
                "required_questions_by_symptom": {
                    "chest_pain": [
                        "pain_location", "pain_quality", "pain_radiation",
                        "associated_symptoms", "onset_timing"
                    ],
                    "headache": [
                        "severity", "onset", "location", "associated_symptoms"
                    ]
                }
            }
        },
        
        # Communication quality
        {
            "type": "communication_quality",
            "config": {
                "check_empathy": True,
                "check_clarity": True,
                "check_safety_netting": True,  # Did agent give return precautions?
                "max_jargon_score": 0.1  # Limit medical jargon
            }
        },
        
        # Escalation appropriateness
        {
            "type": "escalation_appropriateness",
            "config": {
                "check_timing": True,          # Did escalation happen promptly?
                "max_escalation_delay_seconds": 60,
                "check_handoff_quality": True   # Was handoff information complete?
            }
        }
    ]
)
```

## ASR (Speech-to-Text) Evaluation

Evaluate your speech recognition separately:

```python theme={null}
# Log ASR output for evaluation
client.log(
    project="asr-evaluation",
    
    input={
        "audio_url": "s3://bucket/call.wav",
        "audio_duration_seconds": 47.8,
        "audio_quality": {"snr_db": 22, "sample_rate": 16000}
    },
    
    output={
        "transcript": "...",
        "word_timings": [...],
        "confidence_scores": [...],
        "asr_provider": "deepgram",
        "asr_model": "nova-2-medical"
    },
    
    expected={
        "transcript": "...",  # Human-verified transcript
        "medical_terms": ["hypertension", "metformin", "dyspnea"]
    }
)

# ASR-specific evaluators
asr_eval = client.evaluations.create(
    name="ASR Medical Accuracy",
    project="asr-evaluation",
    dataset="ds_transcribed_calls",
    
    evaluators=[
        {
            "type": "wer",  # Word Error Rate
            "config": {
                "normalize": True,
                "ignore_case": True,
                "ignore_punctuation": True
            }
        },
        {
            "type": "medical_term_accuracy",
            "config": {
                "term_categories": ["medications", "conditions", "symptoms", "anatomy"],
                "phonetic_similarity_threshold": 0.8
            }
        },
        {
            "type": "speaker_diarization_accuracy",
            "config": {
                "check_speaker_labels": True,
                "check_turn_boundaries": True
            }
        }
    ]
)
```

## Multimodal Evaluation

For systems that combine voice with other modalities (images, documents):

```python theme={null}
# Log multimodal interaction
client.log(
    project="multimodal-triage",
    
    input={
        # Voice component
        "audio_url": "s3://bucket/call.wav",
        "transcript": [...],
        
        # Image component (patient-submitted photo)
        "image_url": "s3://bucket/rash_photo.jpg",
        "image_description": "Patient submitted photo of arm rash",
        
        # Prior records
        "patient_history": {
            "conditions": ["diabetes"],
            "medications": ["metformin"]
        }
    },
    
    output={
        # Voice analysis
        "voice_symptoms": ["itching", "spreading rash", "fever"],
        "voice_triage": "urgent",
        
        # Image analysis
        "image_findings": {
            "rash_type": "maculopapular",
            "distribution": "arm",
            "severity": "moderate"
        },
        "image_triage": "urgent",
        
        # Combined decision
        "final_triage": "urgent",
        "reasoning": "Maculopapular rash with fever in diabetic patient - needs same-day evaluation"
    }
)

# Multimodal evaluators
multimodal_eval = client.evaluations.create(
    name="Multimodal Triage Eval",
    project="multimodal-triage",
    dataset="ds_multimodal_cases",
    
    evaluators=[
        # Voice-specific evaluation
        {
            "type": "triage_accuracy",
            "config": {"source": "voice_triage"}
        },
        
        # Image-specific evaluation
        {
            "type": "image_finding_accuracy",
            "config": {
                "finding_types": ["rash_type", "distribution", "severity"]
            }
        },
        
        # Integration evaluation - did voice + image combine correctly?
        {
            "type": "multimodal_integration",
            "config": {
                "check_consistency": True,   # Voice and image findings align
                "check_comprehensive": True, # Both modalities considered
                "weight_voice": 0.4,
                "weight_image": 0.6
            }
        }
    ]
)
```

## Critical Metrics for Voice AI

Key metrics to track for healthcare voice systems:

| Metric                   | Target  | Why It Matters                             |
| ------------------------ | ------- | ------------------------------------------ |
| **Triage Accuracy**      | > 95%   | Core safety metric - correct urgency level |
| **Red Flag Recall**      | > 99%   | Must catch nearly all danger signs         |
| **Under-Triage Rate**    | \< 1%   | Missing urgent cases is unacceptable       |
| **Escalation Latency**   | \< 60s  | How quickly urgent cases are escalated     |
| **ASR Medical WER**      | \< 5%   | Transcription accuracy for clinical terms  |
| **Symptom F1**           | > 90%   | Correct symptom extraction                 |
| **Guideline Adherence**  | > 90%   | Following clinical protocols               |
| **Patient Satisfaction** | > 4.0/5 | Empathy, clarity, helpfulness              |

<Tip>
  Set asymmetric thresholds. It's better to over-triage (send someone to ER who didn't need it) than under-triage (send someone home who needed emergency care).
</Tip>

## Setting Up Clinician Review

Route flagged calls to physicians and nurses for expert review:

```python theme={null}
# Configure review routing for voice calls
client.projects.update(
    project="patient-triage-voice",
    
    review_config={
        # Automatic flagging rules
        "auto_flag_rules": [
            # Always review emergency calls
            {"condition": "triage_level == 'emergency'", "priority": "high"},
            
            # Review when AI is uncertain
            {"condition": "confidence < 0.8", "priority": "medium"},
            
            # Review red flag cases
            {"condition": "red_flags_detected.length > 0", "priority": "high"},
            
            # Review potential under-triage
            {"condition": "symptoms.contains('chest_pain') AND triage_level != 'emergency'", "priority": "urgent"}
        ],
        
        # Random sampling for QA
        "sampling_config": {
            "routine_calls": 0.05,      # 5% of routine calls
            "semi_urgent_calls": 0.10,  # 10% of semi-urgent
            "urgent_calls": 0.20        # 20% of urgent
        },
        
        # Reviewer requirements
        "reviewer_requirements": {
            "emergency_calls": ["MD", "DO", "NP"],
            "urgent_calls": ["MD", "DO", "NP", "RN"],
            "routine_calls": ["RN", "LPN"]
        },
        
        # Review interface config
        "review_interface": {
            "show_audio_player": True,
            "show_transcript": True,
            "show_ai_reasoning": True,
            "require_triage_override": True,
            "require_notes": True
        }
    }
)
```

## Integration Example

Complete example integrating Akhara into a voice triage pipeline:

```python theme={null}
from akhara import Akhara
from your_asr import transcribe_audio
from your_nlu import extract_symptoms, classify_intent
from your_triage import determine_triage_level

client = Akhara(api_key="gr_live_xxx")

async def process_triage_call(call_audio_url: str, call_id: str):
    """Process a patient triage call through the full pipeline."""
    
    # Step 1: Transcribe audio
    transcript = await transcribe_audio(call_audio_url)
    
    # Step 2: Extract clinical information
    symptoms = extract_symptoms(transcript)
    intent = classify_intent(transcript)
    
    # Step 3: Determine triage level
    triage_result = determine_triage_level(
        symptoms=symptoms,
        intent=intent,
        transcript=transcript
    )
    
    # Step 4: Log to Akhara for evaluation
    client.calls.log(
        project="patient-triage-voice",
        
        audio_url=call_audio_url,
        
        transcript=[
            {"speaker": turn.speaker, "text": turn.text, "start": turn.start, "end": turn.end}
            for turn in transcript.turns
        ],
        
        ai_decision={
            "extracted_symptoms": [s.to_dict() for s in symptoms],
            "intent": intent.label,
            "red_flags_detected": [rf.name for rf in triage_result.red_flags],
            "triage_level": triage_result.level,
            "recommended_action": triage_result.action,
            "confidence": triage_result.confidence,
            "protocol_used": triage_result.protocol
        },
        
        # Pipeline metadata for debugging
        metadata={
            "call_id": call_id,
            "asr_provider": "deepgram",
            "asr_confidence": transcript.confidence,
            "nlu_model": "clinical-bert-v3",
            "triage_model": "triage-ensemble-v2",
            "pipeline_latency_ms": get_pipeline_latency()
        }
    )
    
    return triage_result

# Run evaluation on logged calls
def run_weekly_evaluation():
    evaluation = client.evaluations.create(
        name=f"Voice Triage - Week of {get_week_start()}",
        project="patient-triage-voice",
        dataset="ds_this_week_calls",
        evaluators=[
            {"type": "triage_accuracy", "config": {...}},
            {"type": "red_flag_detection", "config": {...}},
            {"type": "symptom_extraction", "config": {...}},
            {"type": "guideline_compliance", "config": {...}}
        ]
    )
    
    # Wait and get results
    results = client.evaluations.wait(evaluation.id)
    
    # Alert if metrics derubric
    if results.metrics["under_triage_rate"] > 0.01:
        send_alert("Under-triage rate exceeded 1% threshold!")
    
    return results
```

## Next Steps

<CardGroup cols={2}>
  <Card title="Create Your First Evaluation" icon="check" href="/evaluation/docs/getting-started/first-evaluation">
    Run evaluators on your voice data
  </Card>

  <Card title="Human Review Design" icon="clipboard-list" href="/evaluation/docs/evaluation-framework/human-review-design">
    Set up human-in-the-loop review
  </Card>

  <Card title="Platform Capabilities" icon="file-lines" href="/evaluation/docs/platform-capabilities">
    Supported data modalities and evaluation surfaces
  </Card>

  <Card title="Evaluating LLMs" icon="microchip" href="/evaluation/docs/onboarding/evaluating-llm">
    LLM-specific evaluation
  </Card>
</CardGroup>
