Over time, Trust and Safety teams deal with a problem where the more fraud that’s caught, the more manual review queues grow. Technically, they can remedy this issue by cutting review volume to protect the customer experience, but when they do this, false declines will creep up instead.
Recent industry data suggests this doesn’t need to be a trade-off, as the programs that are closing this gap are instead relying on better decisioning rather than bigger queues.
The hidden cost of manual review
While manual review can serve as a safety net, it’s an expensive one.
According to the 2026 True Cost of Fraud Study for Retail and Ecommerce in North America, the total cost of fraud in the United States now exceeds $5.13 for every $1 lost to fraud directly. This loss is largely driven by labor-intensive work like manual review, chargeback handling, and customer service overhead. That multiplier has climbed for years, and every order routed to a human reviewer only adds to it.
The big problem with manual review is that it costs a business twice: once in direct labor, and again in the time it adds to checkout or account approval. Fraud fighters know the queue backs up fastest during peak season, exactly when false declines and genuine fraud attempts both spike. A team that cannot triage volume quickly ends up approving orders on a deadline rather than on the strength of the signal, which defeats the purpose of review in the first place.
False declines are the bigger revenue leak
Nearly half of merchants estimate that up to 5% of legitimate orders are incorrectly declined as fraudulent, according to a March 2026 PYMNTS Intelligence report, and the industry-wide revenue lost to false declines is now estimated at $50 billion. A related PYMNTS report from the year prior put the figure even higher for U.S. e-commerce alone. An estimated $157 billion in sales were at risk from false declines, $81 billion were projected as permanently lost despite recovery attempts, and 47% of retailers stated that false declines had a severely negative effect on customer satisfaction.
Those numbers should reframe how fraud fighters think about manual review. A queue that stops fraud but also strands good customers is doing imprecise work, no matter how careful its rules look on paper.
What separates high-maturity fraud programs from the rest
The gap shows up clearly by region. Adyen’s 2026 fraud report found that APAC merchants reported the heaviest fraud-prevention burden globally in 2024 and 2025. Around 70% cited rising manual review costs, 60% reported higher false declines, and one in three said they couldn’t resolve the tradeoff between blocking fraud and approving legitimate customers.
Markets with deeper authentication adoption tell a different story. Among merchants in Japan, Australia, and Singapore, Adyen’s platform data shows risk approval rates reaching up to 99.57%, up an average of 17 basis points year over year, with chargeback rates falling consistently across all three markets.
The difference comes down to decisioning quality, meaning how many real signals a system evaluates before a case requires manual review from a human, and how confidently it can act on those signals in real time. Programs that invest in stronger automated decisioning both catch more fraud and create less friction, because the model is doing more of the work that used to fall to a reviewer. This is how mature Trust and Safety excels. Fewer cases get escalated to human workers, and the ones that do are the ones that actually warrant manual review.
Build a decisioning layer that earns the queue’s trust
To cut manual review responsibly, you need to start upstream of the queue instead of inside of it. Here’s how you can build an effective decisioning layer:
- Feed decisioning more signals, not additional rules. For example, a Sift Score is built from thousands of behavioral, device, and payment signals across the user journey gives fraud fighters a real-time read (1 to 100, where 1 is trustworthy and 100 is likely fraud) rather than a static rule tripped by one data point.
- Route only genuine ambiguity to people. Risk-based friction can apply step-up authentication, a one-time passcode, or a document check, only to the transactions that actually need it, instead of pushing everything above a blunt threshold into a queue.
- Automate the clear cases on both ends. Workflows should auto-approve high-confidence legitimate orders and auto-decline high-confidence fraud, leaving reviewers to spend their time on the narrow band in the middle where a human judgment call adds real value.
- Watch approval rate and manual review rate on the same dashboard. Insights that track both metrics side by side make it obvious when a rule tightened for fraud is also quietly tightening the funnel.
Set a 90-day benchmark and prove it out
Benchmarks only help if a team measures itself against them consistently. Start by pulling three numbers from your last two full quarters, including manual review rate as a share of total orders, false decline rate, and confirmed fraud loss rate. Compare false decline rate against the 2% to 10% range reported by Merchant Risk Council and PYMNTS Intelligence surveys, then use those three baseline numbers to set a 90-day improvement target for each, rather than just for fraud loss.
A 90-day window is long enough to see a real shift in decisioning accuracy and short enough to tie results to a specific change, whether that is a new Sift Score threshold, a revised Workflow, or a change to which cases route to Dynamic Friction instead of a human reviewer. Track those three numbers on a weekly basis instead of waiting for a quarterly report. Manual review rate should fall first. False decline rate should hold steady or fall alongside it. Fraud loss rate should not creep up as a side effect. If fraud loss rises while the other two fall, the threshold moved too far, which is useful information in its own right, not a failed pilot.
Measure the trade-off honestly, then close it
Fraud fighters should track false decline rate and manual review rate as a pair. A rule that lowers fraud losses by 2% but raises false declines by 5% has not improved the program, it has simply moved the cost from a chargeback line to a lost-revenue line that is harder to see. Reviewing that pair regularly, alongside review queue volume and reviewer accuracy, is how a team proves that automation is reducing total cost rather than just relocating it. Sift’s Fraud Industry Benchmarking Resource (FIBR) is a useful way to check those numbers against peers in your vertical, so a team isn’t just judging its own trend line in isolation.
If your team is having trouble with balancing rising manual review costs against false declines, Sift might be a good fit for you. Sift’s decisioning layer helps teams cut manual review volume without letting false declines creep up, using a real-time risk score and Workflows built to route only genuine ambiguity to a human reviewer. Try out Sift for yourself by requesting a demo today.
Frequently asked questions
Does reducing manual review always mean approving more risk?
No it doesn’t. The goal is to route more of the clearly legitimate and clearly fraudulent volume to automated decisioning, and reduce the amount of time reviewers spend manually reviewing cases. By reducing manual review, they will have more time to devote their energy to genuinely ambiguous cases. Programs that do this well tend to see both a lower manual review rate and fewer false declines, because the model is more precise than a blanket queue.
What is a healthy false decline rate to benchmark against?
Merchant Risk Council and PYMNTS Intelligence surveys put the industry range between 2% and 10% of orders, with many merchants clustered near the middle of that band. Any number in that range still represents meaningful lost revenue, so the trend matters more than the specific benchmark: is the rate falling as decisioning improves?
How does Sift help reduce false declines without adding fraud risk?
Sift aggregates thousands of signals from across the user journey into a single, real-time read on trustworthiness. Because it updates as new information arrives, a program can act with more confidence on marginal cases instead of defaulting to a decline or a manual queue out of caution.





