Payment agent / regression lab
Same instruction.
Different agent.
Find out what changed before a payment goes live.
Run both versions, inspect the balances, replay the evidence.
Compare payment behavior
Claude chooses the actions. Anvil executes and checks the payment.
Checking Claude configuration…
Configure Claudeclaude-payment / v1-policy
Claude follows the approved payment policy.
Policy-focused promptA shorter payment prompt. The model chooses its own actions.
Version under test0x7099…79C8One transfer · maximum 10 tokensVersion comparison
Regression detectedExecution history
Loading saved reports…
Loading evidence…
The source of truth
Approve the intent.
Then test the action.
The evaluator checks these explicit rules. Changing an instruction or policy invalidates its approval.
Vendor payment
Pay exactly 10 TEST-USD to vendor A on Base. Make one payment.
- Network
- Base simulation / chain ID 8453
- Token
- TEST-USD, 6 decimals. A locally deployed test contract.
- Sender
- 0xf39Fd6e51aad88F6F4ce6aB8827279cffFb92266
- Recipient
- 0x70997970C51812dc3A010C7d01b50e0d17dc79C8
- Expected amount
- 10.000000 tokens / 10000000 raw units
- Maximum total spend
- 10 tokens
- Maximum transfers
- 1
The demo creates an approved copy after you check the approval box. Custom scenarios require the CLI approval command.
Create a scenario for your own agent
npm run cli -- init --out scenario.json # Edit the instruction and policy, then: npm run cli -- approve --scenario scenario.json npm run cli -- run --scenario scenario.json --profile baseline
This MVP executes its own test-token contracts. Payment-policy inputs use decimal strings to avoid floating-point rounding.
From fixture to your agent
Keep your model.
Verify its payments.
Use Claude directly or implement the TypeScript adapter. The same deterministic evaluator checks both.
Run a live Claude trial
Checking configuration…
The workspace owner adds these variables in Coolify, then redeploys. For local development, use .env. Keep the API key on the server.
Return to Evaluation runs, choose Claude API, approve the policy and run a comparison. Both prompts call the model; failures are observed, never forced.
Claude API requires an active Console key and available credits. The public demo shares a daily trial allowance. Reports record the model, prompt hash, token usage and transaction evidence.
ANTHROPIC_API_KEY=your_key CLAUDE_MODEL=your_available_model_id CLAUDE_DAILY_TRIAL_LIMIT=6 npm run cli -- run --scenario scenario.json --adapter claude
Replay recorded actions
A replay uses the saved actions and policy. It makes no model call. A new Claude request is always a new trial.
npm run cli -- replay --report artifacts/REPORT_ID.json
Pin a Base fork
Provide an archive-capable Base RPC and an explicit block. Only the local Anvil node receives transactions.
Without these settings, tests use local fixtures with Base's chain ID. They do not reproduce mainnet state.
BASE_RPC_URL=https://your-base-archive-rpc BASE_FORK_BLOCK=your_fixed_block_number
Bring a custom adapter
Export an AgentAdapter with metadata and run(context). Send transfers through context.transfer to preserve evidence.
The CLI accepts a trusted local JavaScript module. Adapter code runs in the Node.js process.
npm run cli -- run --scenario scenario.json --adapter ./my-agent.js