A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings, was evaluated on 4,023 English-language HealthBench questions (80.5% of the benchmark). Scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, physician-written rubrics, and full response and scoring outputs are public for independent verification. A 500-question subset was re-run against current-generation models to test robustness to newer models and judge lineage.
A corpus-specific clinical RAG system, VITA, designed for low- and middle-income settings, outperformed newer frontier LLMs on the HealthBench medical benchmark. On 4,023 English-language questions, VITA achieved 51.9% of possible rubric points, surpassing GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6. The system uses a curated corpus of India-specific guidelines, antimicrobial resistance data, formulary constraints, and resource-limited care protocols. The benchmark, rubrics, and scoring outputs are public, while VITA's architecture and corpus remain proprietary.
The result suggests that domain-specific retrieval over a carefully curated corpus can outperform general-purpose frontier models on specialized benchmarks, even when judged by a GPT-4.1 model. The public availability of the benchmark and scoring outputs enables independent verification, while the proprietary nature of VITA's architecture limits direct technical replication. The re-run on a 500-question subset against current-generation models indicates an effort to assess robustness to model updates and judge lineage, though details of those results are not provided in the evidence.
This evidence highlights a competitive dynamic where specialized, context-aware systems can challenge general-purpose LLMs in high-stakes domains like healthcare, particularly in low-resource settings. The focus on LMIC-specific data (India-specific antimicrobial resistance, formulary constraints) addresses a gap often overlooked by frontier models trained on predominantly high-income data. The public release of benchmark and scoring outputs may encourage further benchmarking and transparency in clinical AI evaluation.
For healthcare providers in LMICs, a system like VITA could offer more accurate and contextually relevant clinical decision support than general-purpose LLMs, potentially improving patient outcomes and reducing costs. For developers, the evidence suggests a viable business model in building domain-specific RAG systems with curated corpora. However, proprietary architecture may pose barriers to integration and trust without independent validation.
Observable next signals include publication of the 500-question subset results, independent replications of the HealthBench evaluation, and potential adoption or scrutiny of VITA in clinical settings. Further benchmarks comparing RAG systems against frontier models in other specialized domains may emerge. The proprietary nature of VITA may limit widespread adoption unless details are disclosed or licensed.