Veralia

For teams running an AI agent that resolves support cases

Your agent says it's working. Is it?

An agent can hallucinate a resolution as easily as an answer, and some platforms count a customer going silent as a success. Every number you have comes from someone with a stake in how it looks. That is part of why so many companies are spending heavily on AI and cannot say what they got for it.

Veralia is an independent measurement firm. We lock the test in writing before we look, score your real cases against evidence the agent cannot touch, and hand your engineers the failing cases with the exact rule each one broke. You learn what your agent actually does. Your team gets what it needs to tune it.

The dashboard says Resolved. What happened The customer gave up.

By their own published definitions, these can be the same event.

The gap

None of this requires anyone to be acting in bad faith.

The number you see and the number you can prove are not the same number.

In research, evidence is ranked: the higher the tier, the harder it is to fake. Agent metrics have the same structure. Your headline figure is drawn from the bottom of the pyramid. We score at the top.

The hierarchy of evidence for your agent

Schematic
The hierarchy of evidence behind an agent's headline number A pyramid of four bars, the three lower tiers hollow and dashed representing unverifiable claims, the solid apex labeled provable outcomes where scoring happens. From the base upward: every inquiry the agent touched, cases the agent says it closed, outcome on record, provable outcomes. Tier widths are illustrative, not measured. Veralia scores at the apex. WHAT SURVIVES SCRUTINY. WE SCORE HERE. Provable outcomes Outcome on record Cases the agent says it closed Every inquiry the agent touched Each step up, claims that cannot be checked fall away
Every layer below the top is a claim that cannot be independently checked. We score at the summit, on outcomes provable without asking the agent.

Three things follow.

01 / DEFINITIONS

Every vendor defines success differently.

One platform counts a resolution after seventy-two hours of silence, no customer confirmation required. Another counts an assumed resolution when the customer leaves without asking for more help. A third lets its own system classify the outcome. The platforms say so themselves: their own documentation warns these figures can include customers who dropped off or were unhappy.

02 / NAMES

The same number goes by at least nine names.

Resolution rate, automated resolution, automation rate, deflection, containment, and the list keeps going. Nine or more names for what is presented as one thing makes comparison impossible.

03 / DENOMINATORS

Some numbers come with no count behind them.

A rate without a denominator cannot be checked.

What we score against instead

Not the agent's opinion of how it did. We score against evidence that exists outside it: a resolution code, a refund issued, a shipment sent, a human stepping in.

Sooner or later someone asks you to prove the agent is working. Answering that question is the only thing we do.

Founder

Who runs this.

Veralia is run by Aidan Rubio, a published scientist at Stanford Medicine. He started in research at sixteen as an NIH research intern and has held research roles across five labs at Stanford and the University of Pennsylvania.

The discipline here is the one that keeps published science honest: decide what counts as a pass before you look, write the sample down so anyone can redraw it, report the result you got.

We are measurement specialists rather than agent builders, and that is deliberate. The person scoring the test should not be the one selling the fix. Your engineers apply every change; we do not advise on implementation.

Independence

Revenue we have foreclosed, in writing.

Independence is a list of things we do not sell.

Independence here is a design constraint, not a slogan. We narrowed the business to measurement alone and turned down every revenue stream that could lean on a result, so the conditions stay scientific. The list is written down so you can hold us to it.

  • NOTWe do not build agents, sell tooling, or implement changes. Nothing we earn depends on how your result comes out.
  • NOTWe do not fix what we find. Your team applies every fix.
  • NOTWe do not touch your production systems. Historical case data only.
  • NOTWe do not bend the rules after the fact. No protocol changes to improve a score, no dropped cases, no untraceable numbers.
  • NOTWe do not let the agent grade itself. Its confidence score and any vendor label are invisible to scoring.
  • FIXEDThe price is agreed before the work starts, so it cannot move with what we find.

The one interest we do have

Design partners get a reduced rate in exchange for a named case study and a reference call. Case studies are worth more to us when the result is good. We would rather you hear that from us now than notice it later.

Method

This is the part you should argue with.

The rules are locked before we see your data.

Everything that could be bent to flatter a result is written down and sealed with a digital fingerprint before any case is read. Change one character of the plan afterward and the fingerprint no longer matches. Clinical trials run on the same rule, called pre-registration: the test is fixed before the results exist. Changes are allowed; silent changes are not: every amendment is numbered, dated, and printed next to the original.

Before we see anything

  1. Which cases count is defined in writing first. Exclusions are declared and counted, never quietly dropped.

  2. The sample is drawn by recorded rules and a recorded seed. Anyone can re-run the draw and get the same cases.

  3. What counts as a pass is fixed per metric before a single case is read. Nothing is added after we see data.

How scoring works

  1. Four standing checks on every case: did the claimed resolution match the recorded outcome, did it hand off when its own policy required, was every factual claim supported by the record, did it stay inside its declared limits.

  2. Every case is scored more than once, independently. Disagreements are resolved against the written rule, and the disagreement rate is published.

Proof the fixes worked

  1. Part of the sample is sealed from the start and never used to tune anything. After your engineers fix, we re-test on a fresh sample under the same locked plan, and report before, after, and holdout side by side, even where they disagree.

Every number can be taken apart

Each scored case carries case id, sample group, seed, scorer ids, timestamp, and rule version. Dispute any figure and we walk you back to the exact cases behind it.

The engagement

Fixed scope, fixed price, fixed end date.

The Agent Reliability Sprint.

One engagement, three moves. Scope, price, and end date are fixed at signature.

STEP 1

Agree the test, then lock it.

You see the full protocol before we run it.

STEP 2

We score your real cases.

Historical data, never production.

STEP 3

Your team fixes.

We re-test and report whether it worked.

Why the sealed holdout matters

Schematic
How the sealed holdout works across the engagement Two tracks across four stages: lock, baseline, your fixes, re-test. The sample track is scored at baseline and again at re-test. The holdout track stays sealed through the first three stages and is scored once, at re-test. 1 LOCK 2 BASELINE 3 YOUR FIXES 4 RE-TEST SAMPLE HOLDOUT SCORED SCORED SEALED SEALED SEALED SCORED ONCE NOBODY SEES THE HOLDOUT UNTIL YOUR FIXES HAVE SHIPPED
Schematic of the engagement, not a result. The holdout is separated at the very start and never seen by anyone until after your fixes ship. It is the difference between "the cases we showed you got better" and "the agent got better." Baseline, re-test and holdout are reported side by side, and where they disagree the record says so.
  • SCOPEOne deployed agent, one channel, one time window.
  • DATAYour choice before anything moves: redacted export we hold, or resident mode inside your systems, read-only.
  • FIXESYour engineers make every change and state in writing what changed and when.
  • OUTNot included: remediation code, agent development, production access, ongoing monitoring.

What you walk away with

A scored, provenance-stamped reliability record: the test, the results, and the proof in one document you can put in front of anyone. What failed is reported as clearly as what passed.

  • the locked pre-registration and its hash
  • every amendment, numbered and dated
  • the per-case scores with provenance stamps
  • before, after, and holdout figures with intervals
  • the failure catalogue grouped by rule, with counts
  • your own written statement of what you changed
  • the scorer disagreement rate
  • the limitations, including what the sample cannot support
  • everything excluded, and why

Qualification

Thirty minutes. A qualification call, not a pitch.

Five things have to be true on your side.

Two people on the first call: the one who answers for the quality number, and the one who can pull the data.

  1. Enough cases to measure.

    At least 300 qualifying cases.

  2. Real recorded outcomes.

    Not just transcripts, and never the agent grading itself.

  3. A way to get us the data.

    One of the two data modes, owner named.

  4. An engineer with time reserved for the fix round.

  5. Someone who can name who approves the budget.

If all five clear, you get a written scope and draft protocol before anything is scored. If any fail, we say which and what would have to change. Not yet is a good outcome.

Limits

What this does not do.

  • SCOPEA sprint measures what happened in a defined sample over a defined window. It does not predict the future, and it says nothing about systems outside its scope.
  • SCOPEWe do not evaluate the underlying model, your vendor, or your people. Anything outside the scope statement is recorded as out of scope, not tested informally.
  • GATE300 is the minimum number of qualifying cases, checked before we start. It is not the sample size, and it is not a claim about statistical power.
  • RECORDThe record below comes from an automated instrument that forecasts public events. It shows how we handle a measurement. It is not evidence about any AI support agent, and it is not a client result.
  • NEWVeralia is new, and nothing here has been run for a client yet. We would rather you read that on this page than find it out on the call.

Before you consider this

The hostile ones are here on purpose.

Questions you should ask us.

What happens if the result is bad, and who sees it?

You do, first. The record is delivered to you, not published. It reports the misses as plainly as the passes; nothing gets softened or moved to an appendix. Every figure traces back to specific scored cases, so a number you dispute can be taken apart rather than argued about.

You are still paid by the company being measured. How is that independent?

It is the right question, and "we are independent" is not an answer to it. The answer is the method, not the money. There is nothing we can sell you afterwards that depends on how the result comes out. The price is fixed before the work starts, so it cannot move with what we find. The rules and the sample are locked before any case is read, changes are allowed but never silent, and the one interest we do have is printed above rather than left for you to find.

Have you done this before?

Not for a client. Veralia is new, and we would rather say so than dress it up. There is no case study and no result to show you. What exists is the method above, written down in full and open to attack, and the separate record below. If you need a completed engagement before you will consider one, we are not the right firm yet, and that is a reasonable position. Design partner slots are open for exactly this reason.

We already measure this internally.

Then one question matters: what sets your recorded outcome? If the agent grades itself, there is nothing independent to score against, and that is a disqualifier here rather than something we work around. It is also worth checking what your headline number actually counts, because cases with provable outcomes are a smaller set than inquiries. The same goes for the dashboard your vendor gives you: it reports the vendor's definition.

You will find a problem and then sell us the fix.

We cannot. We do not implement, advise on implementation, or touch any system. Fix-it code, agent development, production access and ongoing monitoring are all outside the engagement, permanently.

Who actually scores the cases?

Scoring is done independently more than once, blind to the agent's own confidence and to any vendor-supplied quality label. Disagreements are logged, resolved against the written rule, and the disagreement rate is published in the record. Ask on the call how that is staffed for your engagement and we will tell you plainly.

Will this touch our production systems, or our customers' personal data?

No production access, ever. We work from historical case data. You choose how it is handled before any of it moves, and one of the two modes never moves the data at all. We do not hold raw personal data, and that is an operating rule rather than a preference.

What is this going to cost?

A fixed price, quoted in writing before any work starts, agreed together with the scope and the end date. There is no figure on this page. Ask and you will have one.

Our own record

Provenance, not accuracy.

We test ourselves the same way.

An automated instrument we operate makes forecasts on public questions and scores them against what actually happened. Before we look at any result, the true outcome is pinned down from independent sources. We grade ourselves against the world, never the other way around.

Matched resolutions, cumulative

196 / 196

Matched against pre-registered truth cards pinned before any result was read.

Independent ground-truth facts

approx. 45

Multiple related markets often resolve off a single fact. This line is never published without the one beside it.

Every resolution in the record, one mark each

Every resolution in the record, one mark each Two hundred and thirteen small squares. One hundred and ninety-six solid squares are matched resolutions, sixteen amber squares were graded on the venue’s own outcome under rules declared in advance, and one crossed square was excluded as a documented source conflict.
196 matched
16 graded on the venue’s outcome
1 excluded

As of July 28, 2026: 196 of 196 wire-gradable resolutions matched pre-registered truth cards pinned from independent sources before any result was read. Multiple related markets often resolve off a single fact; the 196 matches rest on roughly 45 independent ground-truth facts.

16 resolutions graded on the venue's own outcome, under rules declared in advance. 1 excluded as a documented source conflict. Both counts cumulative.
The instrument has missed a day. The missed day stands in the record rather than being re-run. Losses, misses and exclusions are kept on the same terms as wins.

Read that for what it is. It proves we lock in the answer before we look and keep our misses on the books. It does not prove our forecasts were good, and we claim no such thing. Discipline is what we sell.

The full record is not published at this address yet. We would rather tell you that than point at a link that does not resolve.

Contact

Design partner slots are open.

The design partner rate is exchanged for a named case study and a reference call.

Tell us a little and we take it from there. Every reply comes from Aidan directly.

If the timing is wrong, say so. Not now is a complete answer.

aidan@veralialabs.com