1. Evidence-first policy assistant

Build retrieval over a small policy collection. Score citation correctness, missing-answer behavior, and document conflicts. Publish the evaluation set.

2. Human-reviewed extraction

Extract structured fields from messy documents, route uncertain cases to review, and measure time saved alongside error rate.

3. Model comparison harness

Compare models on representative tasks across quality, latency, price, and refusal behavior. Explain why the winner changes by use case.

4. Data drift monitor

Train a simple model, simulate a shift in incoming data, and build alerts tied to real performance.

5. Accessible AI interface

Redesign an AI workflow for keyboard, screen-reader, plain-language, and low-confidence use. Test it with users and report what changed.