The challenge
Atlas wanted an assistant but had seen competitors ship chatbots that invented flight times. Anything inaccurate would be worse than no assistant at all.
What we did
- 1Built a retrieval pipeline over live inventory and policy documents
- 2Forced structured outputs so every answer links to a source record
- 3Added a refusal policy and one-click escalation to a human agent
- 4Created an evaluation suite of 400 golden questions run on every deploy
- 5Shipped behind a feature flag to 5% of traffic first
The evaluation harness was the part we did not know we needed. It is why we trusted the launch.
