Evaluate Delvi
Delvi Platform Model Card
Our answers to the five evaluation domains of the AMA AI Specialty Collaborative's AI Tool Evaluation Guide (Feb 2026). Prepared by Delvi; this document does not imply review or endorsement by the AMA.
Evaluating on behalf of a committee? This page prints cleanly — use your browser's Print to save it as a PDF and share it.
Domain 1 · Clinical use case and user
What Delvi is — and is not
- Purpose
- Professional knowledge support for medical society members: literature search and summarization with citations, clinical and administrative writing assistance, research assistance, and self-directed learning over the society's own vetted corpus.
- Intended users
- Licensed medical professionals who are members of the deploying society. Not patient-facing.
- Intended setting
- Professional reference and productivity use. Delvi is not a diagnostic device, does not automate clinical actions, and is not intended for emergency or time-critical clinical decision-making.
- Human oversight
- Every output is advisory. Responses cite their sources so the professional can independently review the basis for any statement before acting on it. Clinical judgment always remains with the clinician.
- Regulatory scope
- Delvi is designed and offered as an informational and reference tool for medical professionals. It is not intended for patient-specific diagnosis or treatment decisions, and every response cites its sources so the professional independently reviews the basis for any statement.
Domain 2 · Training, testing, and validation data relevance
Knowledge sources and grounding
- Grounding corpus
- Responses are grounded via retrieval-augmented generation in the deploying society's own vetted literature — guidelines, journals, clinical algorithms, and educational materials — selected and governed by the society's leadership. The corpus is specialty-specific by construction.
- Citations
- Every substantive response lists the sources consulted, linked to the underlying documents, enabling independent verification.
- Foundation models
- Commercial frontier large language models accessed through a private cloud environment. Delvi does not train models on society content or member data; society content is used at retrieval time only.
- Corpus updates
- New literature is ingested on a cadence set with the society, with document-level metadata (source, publication date) preserved and visible.
Domain 3 · Risks and mitigation
Known failure modes and safeguards
- Confabulation
- Like all generative AI, the underlying models can produce plausible but incorrect statements. Delvi's mitigations are enforced at the system level: responses may only make claims they can cite to the retrieved corpus; when a topic is not covered, the response states that it was not found in the society's library rather than answering from the model's general knowledge; and every claim carries an inline citation for independent verification.
- Scope limits
- Not intended for pediatric, emergency, or any population/setting outside the deploying society's corpus coverage; not a substitute for primary literature review in research publication.
- Oversight & accountability
- The society's leadership defines acceptable-use and best-practice guidance for members; Delvi operates the platform. Outputs are reviewed and acted on solely by the professional user.
- Security & privacy
- Each society runs in a private, dedicated data environment. User conversations are not used to train models. Usage data is aggregated and de-identified for reporting; individual members are not identified to the society.
Domain 4 · Effectiveness and performance
How Delvi is evaluated
- Pre-deployment
- Structured beta evaluation by the society's designated knowledge experts, who test responses against their own literature before member launch.
- Core quality metric
- Citation verifiability — whether responses are supported by the sources they cite. We deliberately do not report diagnostic-accuracy metrics, because diagnosis is outside Delvi's intended use.
- In production
- Member feedback channels and society-level review of aggregate usage inform ongoing evaluation and corpus curation.
Domain 5 · Workflow integration and monitoring
Deployment, monitoring, and updates
- Form factor
- White-labeled web and mobile application under the society's brand. Standalone by design — no EHR integration, keeping Delvi outside clinical workflows and documentation systems.
- Monitoring
- Aggregated usage and trend reporting is provided to society leadership, giving the society direct visibility into how members use AI.
- Updates
- Delvi manages platform and model updates and communicates material changes to the society. Every document in the corpus carries its source, edition, and publication date, and superseded editions are replaced at the society's direction — so members always work from the society's current literature.
- Accountability split
- Delvi is accountable for platform operation, availability, and model behavior; the society governs content selection and member guidance — a shared model documented in each partnership agreement.