You will work on the heart of what we do: the engine that reads an AI-generated clinical note and checks every statement against the source of what actually happened, then proves it or flags it. This is the "umpire" for AI documentation. Your job is to make its calls trustworthy.
What you will do
- Develop and refine the verification engine that compares clinical notes to source records: semantic and token matching, negation and scope handling, numeric and dose-mismatch detection, and flagging of unsupported or fabricated statements.
- Design evaluation methods and test sets that measure what the engine catches and what it misses, with precision and recall that hold up to clinical and audit scrutiny.
- Work with large language models in a grounded, cited, human-in-the-loop pipeline where every output traces back to a source. No source, no claim.
- Help turn model behaviour into reproducible accuracy benchmarks we can stand behind.
Who you are
- A graduate student in computer science, machine learning, computational linguistics, or a related field.
- Comfortable in Python and/or JavaScript, with a working understanding of NLP and LLM concepts and sound evaluation methodology.
- Genuinely interested in safe, verifiable AI, and careful about the difference between a plausible answer and a correct one.
Nice to have
- Experience with clinical or biomedical text, information extraction, or evaluating generative models.