Executive Summary
Agentic commerce is no longer a thesis — it is a deployment. In 2026, personal shopping agents moved from demos into production. ChatGPT shops through connectors and instant checkout; Google's AI Mode summarizes and recommends products directly in search; agentic browsers and smart-home agents reorder household goods without a human opening a product page. Amazon itself now brokers purchases on other retailers' sites through its own agent. The buying side of the internet is being rebuilt around machine readers.
The distribution layer is maturing fast. The Model Context Protocol (MCP), open-sourced by Anthropic in late 2024 and since adopted across every major agent platform, has become the standard rail through which agents discover and call external capabilities — product data, inventory, payments, and increasingly trust. Market analysts size the MCP market at $1.20 billion in 2025, growing to $28.36 billion by 2035 (37.2% CAGR), with AI agents and automation its leading application. GoBuy's own MCP server — live on 7+ channels with 180+ agent installs — is a small but concrete proof point: when agents are given a free choice of tools, they install trust.
But one critical layer is missing: verification. Today's agents are excellent at reading and transacting, and almost universally unequipped to verify what they read. The product web they inherit was built to persuade humans — with review systems Amazon says required blocking more than 275 million suspected fakes in 2024 alone, with pricing that shifts by the session, and with structured data that is missing or wrong on a large share of pages. Without independent evidence scoring, agents do not merely inherit this misinformation — they amplify it, repeating fake signals at machine speed and machine scale.
GoBuy's Evidence Engine has scored [DATA: corpus-size] products across Amazon, Walmart, Target, and Best Buy. The data is blunt: [DATA: pct-insufficient-evidence]% of products lack sufficient marketplace evidence for a careful agent to recommend them; [DATA: pct-missing-schema]% of product pages are missing complete structured data; and [DATA: pct-suspicious-reviews]% exhibit suspicious review patterns. The median product is effectively invisible or unverifiable to the agentic layer.
This whitepaper quantifies the problem and proposes the path forward: a five-layer trust framework — Accessibility, Understanding, Evidence, Transaction, Accountability — that defines what an agent must be able to do before its recommendation deserves confidence, and what brands must fix to stay recommendable. Trust scoring, we argue, is set to become market infrastructure for agentic commerce — the credit bureau of the machine-readable web. The brands that treat their evidence layer as an asset to be managed, rather than a byproduct, will own the agent channel. The rest will be skipped silently.
Part 1 · The Market
Agents Enter Production
1.1 — From search results to agent shortlists
For two decades, e-commerce ran on a simple loop: a human types a query, a search engine ranks results, the human clicks, reads, and decides. Every discipline of the industry — SEO, conversion optimization, review management — optimized that loop for a human reader with human eyes.
That loop is being replaced. In 2026, the query increasingly goes to an agent, and the agent does the reading, comparing, and — in a growing number of flows — the buying. The shift is visible across every major platform:
- OpenAI ships shopping inside ChatGPT via connectors and agentic checkout pilots; its Operator-class agents browse and transact on the open web.
- Google folds product recommendations directly into AI Mode, collapsing the ten blue links into a shortlist the user never inspects page-by-page.
- Amazon — the platform that perfected the human loop — now runs agent-mediated purchasing across other retailers' catalogs, letting its agent complete purchases on Walmart's and Best Buy's behalf-checked storefronts.
- Agentic browsers and assistants (Perplexity-class shopping flows, agentic extensions, smart-home reorder agents) turn routine replenishment into a background process.
- The smart home is becoming a buyer. In September 2026, Google Home began rolling out an MCP server that lets any MCP-speaking agent inspect a household's device states and event history — a running log of consumption that maps directly onto replenishment cycles.
The common thread: the unit of commerce is shifting from the pageview to the recommendation. When a human clicks through ten results, each page gets a chance to persuade. When an agent builds a shortlist of three, seven products never get mentioned. Distribution through agents is winner-take-most.
1.2 — MCP: the distribution rail for agent capabilities
None of this works without a protocol layer. The Model Context Protocol — open-sourced by Anthropic in November 2024 and adopted within months by OpenAI, Google, and the broader ecosystem — standardized how agents discover and invoke external tools. In practice, MCP has become the app store of agentic commerce: a capability published once is installable by any agent, on any platform.
The trajectory is not speculative. Analysts at SNS Insider size the MCP market at $1.20 billion in 2025, heading to $28.36 billion by 2035 — a 37.2% CAGR — with AI agents and automation already the leading application segment at roughly 35% share. Whatever the precise number, the direction is unambiguous: the rail is being laid now, and tools that ride it compound.
GoBuy's MCP server is a live data point on what agents choose when the choice is free. Live across 7+ channels with 180+ agent installs, it offers product Evidence Scores to any agent that asks — and agents ask. Developers building shopping assistants, price trackers, and research agents install a trust-scoring tool not because they must, but because their own output quality depends on not recommending manipulated products. Demand for trust infrastructure is organic and bottom-up.
1.3 — The infrastructure gap: agents can browse, but can they trust?
Stack the agentic commerce layers and the gap becomes visible:
- Reading — solved. Modern agents fetch, render, and parse pages as well as a careful human reader, and better than a careless one.
- Comparing — solved. Structured extraction across multiple listings is a solved retrieval-and-normalization problem.
- Transacting — maturing fast. Agentic checkout, payments APIs, and platform connectors are in production at every major player.
- Verifying — missing. No major agent platform independently verifies that a product's reviews are authentic, its seller is legitimate, or its price signal is stable. Agents treat marketplace data as ground truth. It is not.
This is not a subtle gap. It is the difference between a library and a rumor mill: both contain text; only one has a provenance system. The first three layers make agents fast. Only the fourth makes them safe to listen to.
Agents have been given eyes, hands, and a wallet. They have not been given judgment — and judgment cannot be prompt-engineered out of a product page. — GoBuy Research, 2026
1.4 — Market sizing and why trust is the rate-limiter
Agentic commerce projections vary widely by definition — some analysts frame agentic-driven transaction volume in the hundreds of billions by decade's end [DATA: agentic-gmv-projection — insert chosen GMV projection + source]. We do not need to take any single forecast as gospel to see the structural point:
- The MCP rail alone is projected to grow ~24× by 2035 (SNS Insider).
- Every major consumer platform has shipped agent-mediated shopping or checkout in some form.
- Replenishment (the most predictable, highest-frequency commerce) is being automated first — via smart homes, subscriptions, and reorder agents.
Trust is the rate-limiter because it is the layer with no supply. Reading tools, comparison engines, and checkout rails are abundant and competing. Independent, cross-marketplace evidence scoring — the thing that tells an agent whether to trust the other three — exists almost nowhere at scale. As agent-mediated volume grows, every point of unverifiable inventory becomes either a fraud risk (if agents ignore evidence) or dead inventory (if they respect it). Both outcomes price trust infrastructure into the market.
The remainder of this report quantifies that unverifiable inventory — and what to do about it.
Part 2 · The Evidence Problem
What Agents See — and What They Can't Verify
2.1 — What agents see vs. what humans see
A human shopper evaluating a product page runs a fast, mostly unconscious Bayesian process: the photos look real, the review average is plausible, the seller name rings no alarms, the price feels right. It is a weak verifier — humans are notoriously manipulable — but it is a verifier.
An agent sees something quite different. It sees the DOM: a machine-readable tree in which the product's identity, price, availability, and review corpus are only as legible as the structured data underneath. Three failure modes dominate:
- Rendering dependence. Content injected client-side after load, behind interaction, or blocked by robots directives is invisible to a fetching agent. What the human sees and what the machine reads are different documents.
- Ambiguous identity. Without clean product identifiers (GTIN, MPN, brand + model), an agent cannot confidently join a product across marketplaces, price trackers, or review corpora. It sees three listings; it cannot prove they are one product.
- Unverified signals. The review average, the star rating, the "best seller" badge — all render fine, and none of them prove anything. An agent that treats them as ground truth is trusting the exact surface an incentive exists to manipulate.
The uncomfortable summary: agents are excellent readers of pages and entirely unequipped readers of evidence. They consume signals; they cannot audit them.
2.2 — The state of product evidence across major marketplaces
GoBuy's Evidence Engine has scored [DATA: corpus-size] products across the four dominant US marketplaces, measuring three evidence pillars: review authenticity, seller history, and price stability. The corpus-level picture:
| Marketplace | Products scored | Median Evidence Score | Sufficient evidence* | Suspicious review patterns |
|---|---|---|---|---|
| Amazon | [DATA: n-amazon] | [DATA: median-amazon] | [DATA: pct-sufficient-amazon]% | [DATA: pct-suspicious-amazon]% |
| Walmart | [DATA: n-walmart] | [DATA: median-walmart] | [DATA: pct-sufficient-walmart]% | [DATA: pct-suspicious-walmart]% |
| Target | [DATA: n-target] | [DATA: median-target] | [DATA: pct-sufficient-target]% | [DATA: pct-suspicious-target]% |
| Best Buy | [DATA: n-bestbuy] | [DATA: median-bestbuy] | [DATA: pct-sufficient-bestbuy]% | [DATA: pct-suspicious-bestbuy]% |
Table 1 — Evidence coverage by marketplace, GoBuy Evidence Engine corpus, [DATA: corpus-window]. *"Sufficient" = Evidence Score ≥ 60 (see Appendix A for methodology and band definitions).
Across the corpus, [DATA: pct-insufficient-evidence]% of products fall below the sufficiency threshold. The distribution is not uniform across categories — high-churn categories (consumer electronics accessories, beauty, supplements) show systematically weaker evidence profiles than considered-purchase categories [DATA: weakest/strongest category figures]. Where review velocity is a marketing lever, evidence quality collapses.
2.3 — Fake review patterns: what the Evidence Engine detects
The scale of review manipulation is no longer disputed, least of all by the platforms. Amazon's own Brand Protection reporting describes blocking more than 275 million suspected fake reviews in 2024. The US Federal Trade Commission's rule on consumer reviews and testimonials (effective October 2024) put civil penalties — up to $51,744 per violation — behind the prohibition of fake or AI-fabricated reviews. Regulation confirms the disease; it does not cure the corpus.
GoBuy's Evidence Engine approaches reviews not as a moderation problem (is this individual review fake?) but as a statistical one (does this corpus look organically generated?). Six pattern families carry most of the signal:
- Burst clustering — review volume spiking far above the product's baseline velocity, typically within days of launch or a price promotion. Organic attention ramps; purchased attention arrives in waves.
- Textual homogeneity — n-gram and embedding similarity across supposedly independent reviewers far above organic baselines. Real reviewers describe the same product with startlingly different prose.
- Rating–text divergence — five-star ratings attached to text whose sentiment or specificity does not support them; the signature of incentivized or copied reviews.
- Reviewer concentration — disproportionate weight from accounts with single reviews, no history, or overlapping review graphs across unrelated products from the same seller.
- Timing anomalies — review arrival correlated with negative-event suppression windows (negative reviews arriving are followed by compensating bursts).
- Verified-purchase inflation — verified badges whose purchase paths (refund clusters, micro-price transactions) indicate laundering rather than buying.
In the GoBuy corpus, [DATA: pct-suspicious-reviews]% of products trip at least one suspicious-pattern threshold, and [DATA: pct-suspicious-multi]% trip two or more [DATA: per-pattern breakdown table optional]. The critical point for agentic commerce: an agent consuming a star average consumes the manipulation with it. No prompt, however careful, extracts truth from a corrupted statistic.
2.4 — Schema implementation gaps: the identity layer is broken
Structured data is the agent's native language, and it is in worse shape than most brands assume. Across the corpus:
- [DATA: pct-missing-schema]% of product pages ship no or incomplete schema.org Product markup.
- [DATA: pct-schema-gtin]% carry complete product identifiers (GTIN/EAN/MPN + brand + model) — the minimum for cross-source identity resolution.
- [DATA: pct-price-availability-mismatch]% show structured price or availability fields that mismatch the rendered page — stale, wrong, or deliberately divergent markup.
The third bullet deserves emphasis. Structured data that contradicts the human-facing page is worse than none: it silently corrupts every downstream comparison. An agent comparing prices across three retailers may be comparing numbers that three different CMS updates forgot to sync. The machine-readable web is a second storefront, and for most brands it is unmaintained.
2.5 — The cost of missing evidence: agents skip
What happens when an agent encounters a product it cannot verify? The emerging production behavior is neither random nor hostile — it is conservative. Well-engineered agents, faced with a tie between a verified product and an unverifiable one, recommend the verified one. Faced with instructions to avoid sponsored or unverified listings, they exclude. The agent does not write a complaint; it simply omits.
This makes the evidence gap a revenue line, not an abstraction:
- [DATA: pct-skipped-agents]% of corpus products fall below the bar a conservative agent would apply to a "safe recommendation" shortlist.
- For brands, the mechanism is a silent tax: paid acquisition still delivers the human click, but the agent channel — already compounding at MCP-rail growth rates — routes around them.
- We estimate the annual value of agent recommendations redirected from insufficient-evidence products to verified competitors at [DATA: trust-gap-cost] [DATA: trust-gap-cost method note — model assumptions].
Nobody sees the products the agent chose not to show. That is exactly why the trust gap is so easy to ignore — and so expensive to ignore. — Part 2, The Evidence Problem
The evidence problem is thus not a content-quality issue or an SEO issue. It is an eligibility issue. In the agent channel, evidence is the admission ticket.
Part 3 · The Framework
The 5 Layers of Agent Trust
Trust in agentic commerce is not one property but a stack. An agent's recommendation deserves confidence only if every layer beneath it holds. GoBuy's framework decomposes agent trust into five layers — each independently measurable, each with distinct failure modes, and each with distinct owners (platform, seller, brand, infrastructure). The framework is the organizing principle behind GoBuy's Evidence Scores and the agent-readiness audits discussed in Part 5.
Accessibility
Can the agent read the product at all?
Before evaluation comes retrieval. Accessibility is the binary gate: server-rendered content (or reliably hydrating content), crawlability under standard agent user-agents, robots directives that permit product fetches, and payloads light enough for an agent's budget. A product behind an interactive wall, a broken canonical, or a blanket robots block does not get evaluated — it does not exist.
- Measure
- % of catalog fetchable and parseable by a standard agent profile; time-to-content; canonical integrity.
- GoBuy data
- [DATA: pct-accessible]% of corpus pages fully accessible to agent fetch profiles; the remainder fail on rendering, robots, or payload limits.
- Failure mode
- Invisibility — the agent cannot consider the product, so it never does.
Understanding
Does the product make sense to the machine?
Given readable content, the agent must resolve it into a coherent product: complete schema.org Product markup, correct identifiers (GTIN/MPN/brand/model), accurate price and availability fields, and context that disambiguates variants. Understanding failures are subtler than access failures: the agent sees something, but it may be the wrong something — a stale price, a merged variant, an orphaned listing.
- Measure
- Schema completeness rate; identifier coverage; structured-vs-rendered consistency; variant model coherence.
- GoBuy data
- Only [DATA: pct-schema-gtin]% of corpus pages carry complete identifiers; [DATA: pct-price-availability-mismatch]% show structured/rendered mismatches (Part 2.4).
- Failure mode
- Ambiguity — the agent cannot prove what the product is, so it cannot compare or recommend it.
Evidence
Is there proof behind the claims?
This is the layer the modern web skipped. Evidence asks whether the signals surrounding the product — reviews, seller track record, price history — actually support the listing's claims. GoBuy's Evidence Engine scores three pillars: review authenticity (the six-pattern statistical audit of Part 2.3), seller history (tenure, fulfillment track record, policy behavior), and price stability (variance, phantom discounting, manipulative repricing). Together they produce the 0–100 Evidence Score.
- Measure
- Evidence Score (0–100) with pillar-level sub-scores; corpus-relative percentiles; band assignment (Appendix A).
- GoBuy data
- Corpus median [DATA: median-score]; [DATA: pct-insufficient-evidence]% below the sufficiency threshold; band distribution [DATA: band-distribution].
- Failure mode
- Manipulation — the agent recommends a product whose signals were purchased, not earned.
Transaction
Can the agent act on the recommendation?
A trusted product the agent cannot transact is a dead end. Transaction-layer trust means machine-actionable pricing and availability, structured offers, and checkout access — via APIs, MCP tools, or platform connectors. It also means the offer's terms are stable: no bait pricing, no post-shortlist surcharges, no availability that evaporates at click time. The FTC's 2026 action against undisclosed ad-surcharge auction mechanics on a major marketplace illustrates how transaction-layer opacity is already a live enforcement frontier.
- Measure
- Structured offer coverage; checkout API/connector availability; price-at-click vs. price-at-recommendation consistency.
- GoBuy data
- [DATA: pct-transaction-ready]% of corpus products expose machine-actionable offers; price-consistency delta [DATA: price-consistency].
- Failure mode
- Abandonment — trust earned upstream is destroyed at the moment of action.
Accountability
Can the decision be explained and audited?
The final layer governs the agent side as much as the product side. A recommendation is trustworthy only if it is reconstructable: which evidence was consulted, at what score, from what snapshot, under what decision policy. Accountability is what turns agent output from an oracle into infrastructure — auditable provenance chains that a user, a brand, or a regulator can inspect after the fact. This layer is the least built today and the most consequential tomorrow.
- Measure
- Decision provenance completeness (evidence snapshot refs, tool calls, policy version); audit-trail availability; explainability of the ranking step.
- GoBuy data
- GoBuy's MCP tool responses carry score provenance by design — evidence snapshot, pillar sub-scores, corpus percentile — making every agent call auditable. [DATA: provenance coverage stat if quantified]
- Failure mode
- Opacity — no one, including the agent's builder, can say why a product was recommended.
3.1 — The stack in summary
| Layer | Question | Owner | GoBuy weight* | Dominant failure |
|---|---|---|---|---|
| 1 · Accessibility | Can it be read? | Brand / platform | [DATA: weight-accessibility] | Invisibility |
| 2 · Understanding | Does it parse? | Brand / platform | [DATA: weight-understanding] | Ambiguity |
| 3 · Evidence | Is it proven? | Marketplace / infrastructure | [DATA: weight-evidence] | Manipulation |
| 4 · Transaction | Can it be acted on? | Platform / payments | [DATA: weight-transaction] | Abandonment |
| 5 · Accountability | Can it be audited? | Agent / infrastructure | [DATA: weight-accountability] | Opacity |
Table 2 — The five-layer trust stack. *Indicative weights in GoBuy's agent-readiness composite (Evidence Score itself weights Layer 3 pillars; see Appendix A). [DATA: layer weights]
Two properties of the stack matter for strategy. First, the layers are strictly ordered — failing Layer 1 makes Layers 2–5 irrelevant, which is why "fix the schema" precedes "build the brand" in Part 5's roadmap. Second, no single party owns the stack — brands control Layers 1–2 inputs, marketplaces control much of Layer 3's raw material, platforms own Layer 4, and Layer 5 is shared between agent builders and infrastructure. Agentic commerce's trust problem is a coordination problem, and coordination problems get solved by shared, independent infrastructure — the argument of Part 4.
Part 4 · The Argument
Why Trust Infrastructure Is Essential
4.1 — The credit-bureau analogy
Consumer lending had a trust problem structurally identical to agentic commerce's: a lender meeting a stranger needs to know whether to extend trust, the stranger has every incentive to misrepresent, and individual due diligence does not scale. The solution was not better intuition on either side — it was independent scoring infrastructure. Credit bureaus became the shared, neutral layer that turned an unscalable trust decision into a lookup.
Three properties made that infrastructure work, and all three apply here:
- Independence — the scorer is not a party to the transaction. A marketplace grading its own listings, or an agent grading its own recommendations, has a conflict the market prices in.
- Coverage breadth — scores are only useful if they cover the whole decision space, not one platform's inventory. Cross-marketplace coverage is the product.
- Methodological transparency — the score is auditable and versioned, so it can be contested, regulated, and improved.
Agentic commerce is building exactly this shape: agents (lenders of attention) meeting products (strangers) at machine speed. What is missing is the bureau.
4.2 — Without verification, agents amplify misinformation
It is tempting to assume agents inherit humanity's fraud problem unchanged — the same fake reviews, just read by a machine. That understates the harm in two directions:
- Speed. A human might encounter a manipulated listing once. An agent encountering it re-reads it for every user, every session, every workflow that touches the category — and each reading is a fresh endorsement opportunity for corrupted signals.
- Confidence. Agents present shortlists with fluent authority. A fabricated 4.9-star average, laundered through a recommendation, acquires the credibility of the assistant that repeated it. The manipulation is not merely propagated; it is laundered.
The FTC's fake-review rule and Amazon's 275M-review enforcement are necessary, but they police inputs at marketplace boundaries. Between the marketplace and the agent's output sits nothing — no independent check, no provenance, no responsibility. That gap is where agentic commerce will either earn durable user trust or destroy it industry-wide. One high-profile "AI assistant recommended a scam product" news cycle is worth more to the argument than any whitepaper.
4.3 — The trust chain: agent → platform → product → evidence
Every agent recommendation implies a chain of trust:
Agent → Platform → Product → Evidence
The user trusts the agent; the agent trusts the platform; the platform hosts the product; the product rests on evidence. The chain is only as strong as its last link — and today the last link is unaudited.
Each arrow is currently implicit and unpriced. The agent platforms assume marketplaces verify; marketplaces assume sellers behave; sellers assume their listings speak for themselves; and the evidence — the only link that can be independently checked — is the one nobody in the chain is contracted to audit. Breaking the chain is cheap for bad actors. Mending it after a failure is expensive for everyone else. The economics only balance when the last link is scored by a party with no stake in the outcome.
4.4 — GoBuy's role: the independent evidence layer
GoBuy occupies exactly one position in this stack, deliberately: the independent evidence layer.
- Not a marketplace. GoBuy sells nothing and ranks no storefront; it has no inventory whose success it prefers.
- Not an affiliate. Recommendation outcomes do not monetize for GoBuy, which is what keeps scores honest when a high-scoring product is a low-revenue one.
- Not a platform agent. GoBuy serves every agent equally over MCP — 7+ channels, 180+ installs — because the bureau's value is coverage, not exclusivity.
The Evidence Engine's corpus — [DATA: corpus-size] products, three scored pillars, auditable methodology (Appendix A) — is the bureau's coverage. The MCP distribution is the bureau's delivery rail. Neither is experimental: both run in production today.
4.5 — Why brands should care: recommendation share is revenue
For brands, the practical translation is one sentence: in the agent channel, trust scores are shelf placement.
- Agents shortlist. Products outside the shortlist do not lose the sale on price or features — they are not considered, and the brand never knows.
- Shortlist position will be decided by verifiable evidence, because that is the only input agents can defend to their users. A brand with a genuinely good product and weak evidence loses to a mediocre product with clean evidence — invisibly, at scale.
- The annual value of recommendations redirected across the corpus by the evidence gap is estimated at [DATA: trust-gap-cost] and compounding with agent-channel growth.
The brands that internalize this early treat their evidence layer the way they once treated search ranking: as a durable, compounding asset requiring deliberate investment. That investment program is Part 5.
Part 5 · The Playbook
Recommendations for Brands
The good news buried in this report's data: most evidence failures are fixable, and most competitors have not started fixing them. A brand that closes its trust gaps now is early by construction. The program is three moves and one calendar.
5.1 — Audit your storefront
Start with the binary question: can agents read and understand your catalog at all? GoBuy's Agent Readiness audit (Station, agentic.gobuy.ai) walks a catalog through Layers 1–2 of the framework — fetchability, schema completeness, identifier coverage, structured/rendered consistency — and returns a per-SKU defect list. The output is deliberately unglamorous: a punch list, not a scorecard. Fix the punch list before anything else; Layers 3–5 are theoretical until Layers 1–2 hold.
5.2 — Understand your Evidence Scores
Next, find out what the evidence layer says about you. GoBuy's Evidence Bureau (audit.gobuy.ai) surfaces your products' Evidence Scores across marketplaces — with pillar-level breakdowns of review authenticity, seller history, and price stability, plus corpus percentiles that show where you actually stand against the category. The point is diagnosis: a low score driven by review-pattern anomalies requires a completely different intervention (and timeline) than one driven by price instability or thin seller history.
5.3 — Fix what's broken
Diagnosis without tooling decays into a slide deck. The Agent Optimization Console (gobuy.ai/ao-console/sign-up) turns audit findings into a managed remediation queue — schema defects, identifier gaps, evidence weaknesses — prioritized by agent-channel impact, with progress tracked against the same scores agents consult. The Console is where the 90-day roadmap below actually gets executed.
5.4 — The 90-day agent optimization roadmap
| Phase | Focus | Key actions | Exit criterion |
|---|---|---|---|
| Days 1–30 Baseline |
Layers 1–2 (Accessibility, Understanding) | Run the Station audit; fix robots/rendering blockers; complete Product schema and identifiers on top-SKU pages; reconcile structured vs. rendered price/availability. | [DATA: target]% of top SKUs fully agent-readable with complete identifiers. |
| Days 31–60 Evidence |
Layer 3 (Evidence) | Pull Bureau scores; address review-pattern flags (review acquisition hygiene, burst response); stabilize pricing and kill phantom-discount patterns; document seller-of-record history. | Evidence Score movement on flagged SKUs; [DATA: target] flagged SKUs cleared or in remediation. |
| Days 61–90 Compound |
Layers 4–5 (Transaction, Accountability) | Expose machine-actionable offers where possible; verify price-at-click consistency; wire Console monitoring for regressions; establish quarterly evidence review as an owned process. | Agent-readiness composite stable; regression alerting live; owner named for the evidence layer. |
Table 3 — The 90-day agent optimization roadmap. Targets should be set from the Day-1 audit baseline. [DATA: phase targets]
One caution to close the playbook: do not attempt to game the evidence layer. The Engine's patterns are corpus-relative and adversarially hardened, and the FTC has criminalized the supply side of fake signals. The only durable strategy is the boring one — real evidence, made machine-readable, kept current. That is also, conveniently, the strategy competitors cannot copy quickly.
Appendix
Methodology, Data Sources, and About
A. Methodology — how GoBuy scores products
The Evidence Engine scores marketplace products on a 0–100 scale from three pillars:
- Review authenticity — statistical audit of the review corpus against the six manipulation-pattern families (Part 2.3): burst clustering, textual homogeneity, rating–text divergence, reviewer concentration, timing anomalies, and verified-purchase inflation. Scored corpus-relatively with Wilson-interval bounds, so thin-but-honest corpora are not punished like manipulated ones.
- Seller history — marketplace tenure, fulfillment track record, policy and dispute behavior, and catalog breadth/consistency of the seller of record.
- Price stability — variance analysis of the price series, phantom-discount detection (anchor inflation), and cross-source price consistency.
Pillar sub-scores are normalized, Wilson-bounded, and combined by a weighted composite — weights [DATA: pillar weights] — into the Evidence Score. Scores are computed per product per marketplace, versioned with the engine, and served with provenance (snapshot reference, sub-scores, percentile) through the MCP tool.
Score bands:
- 0–39 · Insufficient — evidence does not support a confident recommendation.
- 40–59 · Weak — partial evidence; recommendable only with caveats.
- 60–79 · Established — sufficient evidence for standard agent recommendation.
- 80–100 · Exemplary — strong, consistent evidence across all pillars.
Limitations. The corpus covers the four dominant US marketplaces and is snapshot-based; scores lag real-world changes by the refresh cadence [DATA: refresh cadence]. Statistical pattern detection identifies manipulation signatures, not legal guilt — scores are decision inputs, not adjudications.
B. Data sources
- GoBuy Evidence Engine corpus: [DATA: corpus-size] products across Amazon, Walmart, Target, and Best Buy; collection window [DATA: corpus-window].
- GoBuy MCP distribution telemetry: 7+ channels, 180+ agent installs (production, as of September 2026).
- Amazon Brand Protection reporting (fake-review enforcement volumes, 2024).
- FTC Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465 (effective October 2024); FTC enforcement actions, 2025–2026.
- SNS Insider, MCP market sizing ($1.20B 2025 → $28.36B 2035, 37.22% CAGR).
- Platform documentation and public reporting on agent commerce launches (OpenAI, Google AI Mode / Google Home MCP, Amazon agentic checkout), 2024–2026. [DATA: add specific citations Replit wants linked]
C. About GoBuy
GoBuy is the trust layer for agentic commerce — an independent evidence-scoring infrastructure for marketplace products. Its Evidence Engine scores products on a 0–100 scale across review authenticity, seller history, and price stability; its MCP server distributes those scores to any AI agent on 7+ channels (180+ installs); and its brand tools — Station (agent-readiness audits), Bureau (Evidence Scores), and the Agent Optimization Console — help brands earn and maintain agent recommendations. GoBuy sells no products, ranks no storefronts, and takes no affiliate positions: independence is the product. Learn more at gobuy.ai and blog.gobuy.ai.
References
- Amazon, Brand Protection Report — reporting on 275M+ suspected fake reviews blocked in 2024. [DATA: link to exact report page]
- Federal Trade Commission, Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, effective October 2024.
- SNS Insider, Model Context Protocol (MCP) Market size and forecast, 2025–2035. [DATA: link]
- Google, Home MCP developer documentation, September 2026.
- Anthropic, Model Context Protocol specification, November 2024 and subsequent revisions.
- GoBuy Evidence Engine methodology, this report, Appendix A. [DATA: methodology page link if public]