Payment fraud detection fell with newer AI in Coinbase test

Bitbuy
fiverr


Coinbase reported Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, despite an unchanged decision policy. The findings challenge the assumption that upgrading a model improves an existing payment screener.

The company’s evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The cohort covered nine weeks before its risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic.

Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions. This isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.

Related Reading

Coinbase says it cut a 90-case AI support test from 1–2 weeks to 30–45 minutes

Results from a fixed historical replay

Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Recall measures the share of fraud cases a model catches; dollar-weighted recall measures how much of the total fraud value it catches.

Sonnet’s recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus’s recall declined 0.8 points. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.

GPT showed why one improving score can be misleading. Its precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. Its fraud flags were more accurate, while more fraud cases and value escaped detection in the replay.

Coinbase's historical replay comparing GPT-5.4 with GPT-5.6 (sol): precision rose 11.5 percentage points, recall fell 20.7 points and dollar-weighted recall fell 21.8 points under a fixed decision policy; these are not live customer losses.Coinbase's historical replay comparing GPT-5.4 with GPT-5.6 (sol): precision rose 11.5 percentage points, recall fell 20.7 points and dollar-weighted recall fell 21.8 points under a fixed decision policy; these are not live customer losses.

The replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.

Coinbase’s earlier online experiment compared adding selective LLM review with the existing models and rules alone. That agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value; it did not compare newer model versions.

Related Reading

Coinbase traced $1.1 million crypto trail behind AI phishing service EvilTokens

In their limitations, the SR-Fraud researchers say the proprietary dataset cannot be released, restricting independent replication and generalization. Their related payment-fraud study first appeared Sept. 23 and was revised Sept. 30, before the October blogs.

A separate case for a custom model

In its Oct. 8 disclosure, Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 percentage points and dollar-weighted recall rose 35.4 points. The company specialized it using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples.

Separately, production measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative reduction. Faster inference and stronger benchmark detection came from different evaluations.

For payment providers, the upgrade question is whether a candidate improves fraud coverage under their actual decision setup. Coinbase recommends testing that configuration first, then evaluating changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.

Related Reading

AI was supposed to take scammers’ jobs, but it gave them superpowers instead



Source link

Bybit

Be the first to comment

Leave a Reply

Your email address will not be published.


*