Agent Foundry / Evals
Test the Whole Agent
Measure whether the full agent finishes real work safely, not just whether one answer sounds good.
Test the model, tools, context, policy, and final outcome as one working system.
What this page delivers
A full task, tested end to end
- 1Task set
- 2Live-like tools
- 3Policy judge
Result
READY · 92 / 100
Emu on the job
Original mechanism
A full task, tested end to end
Follow the work from input to a result your team can review.
Task set
Live-like tools
Policy judge
Eval receipt
The question answered
How do we know an agent is ready for real work?
Give it representative tasks, realistic tool states, and clear success rules. Then review the full trace, including what it did, what it avoided, and why it passed.
Three-step flow
From a clear input to a governed result.
Build the task set
Mix normal work, hard edge cases, policy traps, and known past failures.
Run the real system
Test the model with its tools, context, identity, and runtime rules attached.
Review the whole trace
See task success, safety checks, tool behavior, costs, and evidence in one receipt.
Capabilities
The controls teams need to go deep.
End-to-end scenarios
Test full jobs with several turns, tools, and decisions.
Multiple judges
Combine rule checks, expected facts, policy results, and human review.
Version comparison
Compare candidate and current agents on the same task set.
Replayable receipts
Keep the setup, trace, scores, and version details for review.
Concrete product artifact
See the proof, not just a green light.
A realistic record keeps the result, the evidence behind it, and the next action in one place.
EVAL RECEIPT
travel-agent / release-27
- Task success
- 46 / 50
- Policy checks
- 150 / 150
- Tool accuracy
- 98.2%
- Critical failures
- 0
› suite: launch-readiness-v9
› agent: travel-agent@27
› receipt: EV-2026-00841
Buyer outcomes
Less guesswork. More control.
System tested
Measure more than the model response.
Work sampled
Use tasks that match what users ask.
For review
Keep the trace behind every score.
Why Trinitite
Built for the full agent lifecycle.
The agent is the test target
Models, tools, context, policy, and side effects are judged together.
A score is not enough
Each result links back to the task, trace, checks, and version that made it.
Failures become reusable assets
A miss can become a unit test, research case, or release blocker.
Keep exploring
Connect this capability to the next handoff.
FAQ
Test the Whole Agent, answered.
What parts of an agent can an eval test?
An eval can test the model response, context, tool calls, policy results, final outcome, and the full path between them.
Can we use our own tasks?
Yes. Real tasks, reviewed production traces, edge cases, and known failures make the strongest task set.
How do we review a score?
Open the eval receipt to see the setup, agent version, full trace, checks, and result behind the score.
Can eval results block a release?
Yes. Release control can require chosen suites and thresholds before a version moves forward.
Build with Agent Foundry
Test the agent your users will meet.
Bring a real workflow. We will show how to turn it into an end-to-end readiness suite.