Context & Opportunity
To explain why a payment is stuck, an Ops analyst rebuilds the case by hand across three disconnected systems, the card processor, the provider that moves the money, and a legacy database linking them, with no single screen that shows all three.
This is agentic work: multi-source retrieval and reasoning. But because money is moving, confidence signals, citations, approval gates, and an audit trail need to be built in from the start.
The flow
A single scenario, a stuck Brightwater Logistics settlement, walked through all three levels and all four transitions. This is the flow the case study is built to demonstrate.
tap the phone to pause · tap a number to jump
Direct Ops data entry
New buyer mapping stays manual until volume justifies the calibration cost.
Compliance & legal approvals
Human, non-negotiable. The agent can assemble the packet; it cannot sign off.
Terminal states of money already moved
Audit yes, action no. Once the money has settled there is nothing left to edit, a correction is a new transaction.
The framework
Three levels of AI sit on one continuum. As you move up, four things shift together: who takes initiative, whether AI can change state, how much trust the interaction requires, and how reversible the output is. What makes it a ladder, not three separate products, is the transition between levels, allowing analysts to move up or down without losing context.
Explain-this-status on a badge, proactive anomaly flags on odd rows, smart empty states, inline summaries inside the drawer that already exists.
Every answer shows its plan, the systems it checked and when, and how sure it is.
An action card with a plan you can preview, approve in bulk or one by one, stop mid-run, and undo. Every step is logged.
The agent flagged it.
Leg 2 was never initiated here or on 11 others, totaling $284,900.
Inline, with the reasoning.
Confidence is stated in words, never a bare number. At L1, the Anomaly flag opens its evidence: $10,000 vs ~$3,100 typical · flagged by Gen 3. At L2, Confidence: medium with the reasoning behind it. At L3 it limits the action: unable to verify one of the 12 transactions, the agent retries 11 and says which one it left out.
Every agent claim is tied to its source, Gen 3 / Transfermate / Gen 2, with the time it was queried. Hover at L1, inline at L2, in the audit trail at L3.
Actions are typed. Safe (retrieval, drafts) needs no approval. Reversible (label a transaction, save a filter) needs a simple approval. Destructive (retry settlement, send external comms) needs approval with details.
If confidence drops at L3, the agent proposes L2, "I'm not sure, want me to investigate?" If L2 can't answer, it proposes a human handoff rather than guessing.
When the agent makes a claim, the reason is always one tap away. Flags show why, answers show their sources, and actions keep Review each, Cancel, and Undo. The agent is never the only way to access the information or take action.
Batch approval for destructive actions
Far faster for the 12-transaction case, but one "approve all" click carries more weight. Mitigated with per-item opt-out and undo, still open whether high-value transactions should force per-item approval.
Open, whose name is on the action?
Accountability blurs between the human who approved and the agent that executed. The audit trail captures the full chain, but the org still has to decide where responsibility sits.
The framework and the prototype behind it are written up on my Medium.