Trinitite joins TechCrunch Startup Battlefield 200

Agent Foundry / Evals

Test the Whole Agent

Measure whether the full agent finishes real work safely, not just whether one answer sounds good.

Test the model, tools, context, policy, and final outcome as one working system.

TRIAL ARENAONLINE

What this page delivers

A full task, tested end to end

  1. 1Task set
  2. 2Live-like tools
  3. 3Policy judge

Result

READY · 92 / 100

Emu on the job

Original mechanism

A full task, tested end to end

Follow the work from input to a result your team can review.

01

Task set

02

Live-like tools

03

Policy judge

04

Eval receipt

arena pathREADY · 92 / 100

The question answered

How do we know an agent is ready for real work?

Give it representative tasks, realistic tool states, and clear success rules. Then review the full trace, including what it did, what it avoided, and why it passed.

Three-step flow

From a clear input to a governed result.

01

Build the task set

Mix normal work, hard edge cases, policy traps, and known past failures.

02

Run the real system

Test the model with its tools, context, identity, and runtime rules attached.

03

Review the whole trace

See task success, safety checks, tool behavior, costs, and evidence in one receipt.

Capabilities

The controls teams need to go deep.

End-to-end scenarios

Test full jobs with several turns, tools, and decisions.

Multiple judges

Combine rule checks, expected facts, policy results, and human review.

Version comparison

Compare candidate and current agents on the same task set.

Replayable receipts

Keep the setup, trace, scores, and version details for review.

Concrete product artifact

See the proof, not just a green light.

A realistic record keeps the result, the evidence behind it, and the next action in one place.

EVAL RECEIPT

travel-agent / release-27

REVIEW READY
Task success
46 / 50
Policy checks
150 / 150
Tool accuracy
98.2%
Critical failures
0

suite: launch-readiness-v9

agent: travel-agent@27

receipt: EV-2026-00841

Buyer outcomes

Less guesswork. More control.

Whole

System tested

Measure more than the model response.

Real

Work sampled

Use tasks that match what users ask.

Proof

For review

Keep the trace behind every score.

Why Trinitite

Built for the full agent lifecycle.

The agent is the test target

Models, tools, context, policy, and side effects are judged together.

A score is not enough

Each result links back to the task, trace, checks, and version that made it.

Failures become reusable assets

A miss can become a unit test, research case, or release blocker.

FAQ

Test the Whole Agent, answered.

  • What parts of an agent can an eval test?

    An eval can test the model response, context, tool calls, policy results, final outcome, and the full path between them.

  • Can we use our own tasks?

    Yes. Real tasks, reviewed production traces, edge cases, and known failures make the strongest task set.

  • How do we review a score?

    Open the eval receipt to see the setup, agent version, full trace, checks, and result behind the score.

  • Can eval results block a release?

    Yes. Release control can require chosen suites and thresholds before a version moves forward.

Build with Agent Foundry

Test the agent your users will meet.

Bring a real workflow. We will show how to turn it into an end-to-end readiness suite.