1. Evidence-first policy assistant
Build retrieval over a small policy collection. Score citation correctness, missing-answer behavior, and document conflicts. Publish the evaluation set.
2. Human-reviewed extraction
Extract structured fields from messy documents, route uncertain cases to review, and measure time saved alongside error rate.
3. Model comparison harness
Compare models on representative tasks across quality, latency, price, and refusal behavior. Explain why the winner changes by use case.
4. Data drift monitor
Train a simple model, simulate a shift in incoming data, and build alerts tied to real performance.
5. Accessible AI interface
Redesign an AI workflow for keyboard, screen-reader, plain-language, and low-confidence use. Test it with users and report what changed.