Agent-to-agent commerce is the practice of software agents buying and selling discrete jobs from each other: summarize this packet, label these images, fetch a priced data slice, or run a bounded tool chain. A small paid task market keeps each job explicit about price, delivery proof, and what happens when results disappoint. This article outlines practical patterns for building such a market without hype and without pretending the hard parts of payments and trust disappear.
What a small paid task market is
A task market lists jobs with a clear input schema, an output schema, a price or price rule, a deadline, and acceptance criteria. Buyer agents post jobs or browse catalogs. Seller agents claim work they can finish inside the stated constraints. Settlement happens when delivery meets the criteria or when a dispute path resolves the case. The market can be a single coordinated service or a thin set of HTTP conventions across independent hosts.
Small is intentional. Early markets work better with a handful of job types, one or two payment rails, and strict size limits on payloads. Breadth can grow after repeatable delivery and refund behavior exist. A sprawling catalog with vague prompts recreates freelancing chaos inside automation and burns budget on retries.
Humans still set policy: which sellers are trusted, maximum spend per hour, and which job categories are allowed. Agents execute inside those rails. The market is infrastructure for priced coordination, not a substitute for organizational judgment.
Pricing patterns that agents can evaluate
Fixed prices per job type keep buyer logic simple. A seller publishes “caption up to N images for amount X” with hard caps on resolution and format. The buyer compares X to its internal value estimate and either pays or skips. Fixed prices pair well with HTTP 402 challenges and prepaid credits.
Tiered prices help when input size varies. Charge by token count, megabyte, or row count with a published formula and a ceiling. Agents must see the quote before funds move. Open-ended “we will bill what it takes” pricing is a poor fit for autonomous buyers that need deterministic spend caps.
Auctions and dynamic bids are optional later features. They add latency and strategic complexity. Most internal or partner markets never need them. If you add bidding, freeze the winner price in a challenge identifier so settlement cannot drift after the buyer commits.
Quote freshness and fees
Quotes should expire. Stale quotes cause sellers to lose money when input costs move, or buyers to overpay when sellers would have discounted. Include the quote expiry in the job offer and in any 402 challenge. State whether marketplace fees sit on top of the seller price or come out of it, and show the all-in amount to the buyer agent.
Currency choice matters for automation. Stable units reduce quote churn. If you settle on chain, document network fees separately so agents do not confuse gas with the job price. If you settle on a ledger credit inside the market, publish the top-up process and the minimum balance sellers and buyers must hold.
Delivery proof without theater
Delivery proof is evidence the seller produced outputs that match the schema and acceptance checks. At minimum the market stores the job id, input hash, output hash or content-addressed pointer, timestamps, and the seller identity that claimed the job. Buyers retrieve outputs through the market or a signed URL after payment clears.
Automatic checks catch many failures: JSON schema validation, row counts, image dimensions, forbidden string patterns, and embedding-similarity floors against a reference. Checks should be published beside the job type so sellers test before claiming. Surprises at acceptance time create avoidable disputes.
Human review remains a backstop for subjective quality. Route only flagged jobs to a reviewer queue. Do not require a person on every five-cent task or the economics collapse. For subjective work, define rubrics with examples of accept and reject so reviewers and seller agents share expectations.
Content addressing and receipts
Hashing inputs and outputs makes later arguments concrete. If both parties can re-hash the bytes, debates shift from “I sent it” to “these bytes failed check twelve.” Storing receipts in an ordinary database is enough for most markets. Public content-addressed storage is optional for open audits or cross-organization citation.
Signed receipts from the market operator help buyers expense the job and help sellers prove completion to their own accounting. Include job id, amount, parties, and output digest. Keep signatures boring and well documented so agent runtimes can verify them in one library call.
Simple dispute paths
Disputes happen when automatic checks pass but the buyer still rejects the work, when checks fail after payment, or when a seller misses the deadline. A simple path is enough to start: buyer opens a dispute with a reason code within a short window; funds stay in escrow; seller may upload one revision; if still unresolved, a market moderator or pre-agreed arbiter decides release, refund, or split.
Reason codes should be a closed list: schema_fail, deadline_miss, policy_violation, quality_fail, duplicate_delivery, other. Free-text alone is hard for agents to branch on. Attach the failing check id when a validator fired. Time-box every stage so escrow cannot sit forever.
Abuse controls matter. Buyers that dispute at extreme rates lose privileges. Sellers that miss deadlines repeatedly lose ranking or access. Publish the rules in machine-readable form as well as prose so agents can simulate risk before claiming a job.
Escrow, payment rails, and agent wallets
Escrow holds buyer funds until acceptance or dispute resolution. The market operator or a specialized facilitator can hold escrow. Agent wallets need scoped keys, spend ceilings, and allowlists of market contracts or HTTP endpoints. Keys with unlimited drain authority have no place in production agents.
Payment can ride x402-style HTTP challenges, prepaid balances inside the market, or invoices settled in batches for trusted pairs. Pick one primary path for the small market and document it. Multiple rails are fine later when volume justifies the engineering.
Refunds should be first-class API operations with the same idempotency care as charges. Agents retry. Double refunds and double captures are classic failure modes when handlers are not idempotent.
Catalog design and discovery
Each job type needs a stable identifier, version, input schema, output schema, price rule, max runtime, and required seller capabilities. Buyer agents search by capability tags rather than by natural-language guesswork alone. Sellers advertise capabilities they actually support in continuous integration tests.
Versioning prevents silent breaks. When a schema changes, bump the job type version and leave the old version available until traffic drains. Agents pin versions in their plans the same way they pin package versions in software builds.
Discovery feeds can be a simple HTTPS index.json updated on a schedule. Fancy semantic search helps only after the catalog is clean. Garbage descriptions with strong search still yield garbage matches.
Operational metrics worth watching
Track time-to-claim, time-to-deliver, automatic accept rate, dispute rate by job type, refund volume, and seller concentration. Concentration risk appears when one seller finishes most jobs and then goes offline. Encourage a second seller for each critical job type before you depend on the market in production paths.
Budget burn per buyer agent should alert humans when spend spikes. Autonomy without spend telemetry is how quiet failures become expensive failures. Pair economic alerts with technical latency alerts so you see both stalled jobs and runaway success that costs too much.
Trust onboarding for seller agents
New seller agents should pass a dry-run suite before they can claim paid jobs. The suite posts sample inputs, checks outputs against validators, and measures latency. Sellers that pass receive a capability badge in the catalog. Buyers can filter for badges when the job is sensitive or when spend is high.
Identity can start simple: an API key bound to an organization, a registered callback URL, and a payout destination. Stronger setups add signed agent attestations or hardware-backed keys. Match identity strength to payout size. Micropayment sellers do not need the same bar as agents that receive large escrow releases.
Offboarding must exist too. When a seller key is compromised, the market revokes claim rights quickly and freezes unsettled escrow for review. Document the revocation API so buyer agents can refresh trust lists on a schedule.
Putting the pieces together in a pilot
A sensible pilot picks three job types, one payment rail, one escrow holder, and two seller teams. Run synthetic buyers for a week with tiny prices. Verify that quotes expire correctly, proofs land once, disputes time out, and metrics dashboards show accept rates. Only then connect a real production buyer agent with a low spend ceiling.
Write the pilot retrospective in prose humans read and in a machine-readable config of what stayed. Promote the config into your market's default templates. Markets improve when operational lessons become defaults rather than folklore in chat threads.
Conclusion
A small paid task market gives agents a bounded way to buy and sell discrete work: clear job types, prices agents can evaluate, delivery proof tied to schemas and hashes, and a short escrowed dispute path when checks disagree. Start narrow, keep quotes fresh, instrument accept and dispute rates, and wrap wallets in hard spend limits. Those patterns are enough to run useful agent-to-agent commerce while ordinary databases, moderators, and payment partners handle the parts that still need durable institutions.



