How do you know if the thing you’re building with AI actually works correctly?

It has never been easier to build software that looks like it works. AI-built apps arrive with a real interface and confident output, looking finished whether or not they are. For a prototype, that’s fine. But some work has to be right: payments that have to reconcile, metadata held to a professional standard, results someone signs their name to before they ship or bill. For that work, “it demos well” is where the trouble starts.

Your experts are the scarce ingredient. We built the system around that.

AI can now generate work at almost any scale: drafts, matches, flags, classifications. What it cannot generate is the judgment that says which of those are right. That judgment lives in the people who do the work today, and it is the scarcest, most valuable input in the whole system.

Human-guided learning is our method for using it well. Your experts don’t write requirements or attend workshops. They do their job inside a working app, and every accept, fix, and rejection they make becomes ground truth, the signal that steers everything the machines do next. The system doing the proposing doesn’t have to be ours, either: if you already run matching or classification software, its output becomes the proposals, and the loop gives it the measured accuracy record it never had.

The agents spend the tokens. Your experts spend the judgment. We make every unit of judgment teach the system as much as possible.

The discovery loop

We don’t start with a spec. We start with a seat next to the people doing the work.

  1. 01

    Sit in.

    We join the people running the manual workflow and learn the job as they actually do it: the inputs, the calls they make, the exceptions that never made it into the documentation.

  2. 02

    Measure the baseline.

    Before improving anything, we instrument the workflow as it stands: volume, time, error patterns. Every claim we make later is relative to this number.

  3. 03

    Ship the first improvement.

    Within days, not months: a working app that takes on part of the workflow and proposes answers. Your experts judge every proposal with one click: agree, fix, or reject.

  4. 04

    Learn from every verdict.

    Each judgment is captured as ground truth. Fixes are the curriculum: they show exactly where the system's proposals miss and why.

  5. 05

    Improve daily.

    Overnight, the system turns yesterday's verdicts into today's version, measurably better, with the change log to prove it.

  6. 06

    Find the frontier.

    Improvement eventually flattens. That's the finding, not the failure. The plateau shows precisely how much of the workflow machines can take at current capability, and how much genuinely requires your people.

The discovery loop. Schematic: the shape of the method, not one particular build of it.

The output isn’t just a better workflow. It’s a measured map of where automation ends and your experts begin, with the evidence to defend it.

Inside the app: self-validating by construction

The app your experts judge every day is built to measure its own accuracy against their judgment, continuously. Nothing runs on autopilot until it has earned it.

  1. 01

    Work arrives

    A claim, a file, a record: whatever your team already handles a hundred times a week.

  2. 02

    The app proposes an answer

    And, just as importantly, how sure it is.

  3. 03

    The gate

    Has it proven, on this kind of work and against the record, that it clears the accuracy bar?

  4. 04

    Yes: handled automatically

    A sample keeps going to a person anyway, so the measurement never goes stale.

  5. 05

    No: routed to a person

    The expert decides, the way they always did.

  6. 06

    Ground truth

    Every call a person makes is kept. That record is the only thing accuracy is measured against.

  7. 07

    Measured accuracy, and a change gate

    The bar moves with the evidence, and no change to the app ships until it has been tested against the record.

The anatomy of a self-validating app. Schematic: it shows the shape of the loop, not one particular build of it.

The old rule was measure twice, cut once: all that care up front because the cut was expensive. AI made cutting free, and a lot of the industry responded by dropping the measuring along with the caution. We went the other way. When cuts are free, measurement is the discipline that’s left: our systems capture the decisions your experts already make and turn them into ground truth, automate only what they can prove they handle at a measured accuracy, route the rest to a person, and never change their own behavior without testing the change against that record first. The accuracy isn’t asserted. It’s measured, and you can check it.

What happens between 6pm and 9am

The daily improvement isn’t a developer reading feedback tickets. It’s an automated development process with your experts’ judgment at the top of it.

When the day’s verdicts land, the system goes to work. Disagreements are analyzed: was the miss a rule the system didn’t know, context it couldn’t see, or a case where reasonable experts would differ? Proposed improvements are drafted, implemented, and then tested against the full history of human judgments before anything ships. A change that fixes today’s misses but breaks last week’s agreements doesn’t ship. Nothing reaches your team without passing the accuracy gate: the new version must beat the old one on the record, not in a demo.

By morning there’s a new version, a plain-language changelog naming which of your corrections drove each change, and fresh numbers on the scoreboard.

No change ships on a hunch. Every version has to beat the last one against your own experts’ judgment.

6PM9AMONE TURN= ONE DAY
6pm → 9am: the agents build, gated by the record9am → 6pm: your experts judge

A small human team, and the agents that never sleep

Signal Foundry runs this process with a deliberately small human team directing a crew of specialized AI agents. Each has one job, and one agent’s output is another agent’s input, with human judgment as the final authority throughout.

Builder agents

Implement proposed improvements: new rules, better context, refined handling for the item types your experts kept correcting.

Evaluator agents

Score every candidate change against the accumulated judgment record before it can ship, and re-score old cases so regressions are caught, not discovered.

The orchestrator

Decides what's worth attempting next. Its objective is judgment-efficiency: of everything we could try, what will learn the most from the fewest demands on your experts' attention?

Monitor agents

Watch the live system for accuracy drift, unusual items, and data-quality problems, and raise their hand when something needs a human look. They're also the ones that call the plateau: when improvement flattens, they flag that the workflow is approaching its frontier.

The reporter

Keeps the record straight: changelogs in plain language, version-over-version metrics, and the frontier report that maps what's machine-ready, what runs assisted, and what stays human territory.

Your experts are not a step in this pipeline; they are its ground truth. The entire agent layer exists to maximize what the system learns from each judgment they make. And on our side, a human reviews every accuracy gate. Agents propose, measure, and flag, but the decision to trust a system with more of your workflow is made by people, on evidence.

More compute doesn’t replace your people. It promotes them.

Here’s the dynamic most AI projects get backwards. Scaling AI is now mostly a matter of spending tokens; the models will happily generate more attempts, more drafts, more candidate answers. But every one of those attempts is worthless until someone qualified says whether it’s right. The real constraint isn’t compute; it’s validated judgment.

That means the value of your experienced people goes up as you automate, not down. They stop doing the repetitive 80% and become the judges of the hard 20%, the ground truth that lets the machines take on more. Human judgment unlocks more automation; more automation makes each judgment more valuable. Our engagements are built to make that flywheel spin, and to measure exactly where it stops.

Where the loop ends: the frontier

Improvement eventually flattens. That’s the finding, not the failure. The plateau shows precisely how much of the workflow machines can take at current capability, and how much genuinely requires your people.

MACHINE-READYASSISTEDHUMAN TERRITORYHANDLED AUTOMATICALLY, SPOT-CHECKEDJUDGED BY A PERSONYOUR EXPERTS
Schematic

If your team has a workflow where somebody answers “is this right?” a hundred times a week, that’s the shape of problem we want. See what an engagement looks like, or write to hello@signalfoundry.ai.