Over 10 years we helping companies reach their financial and branding goals. Onum is a values-driven SEO agency dedicated.

LATEST NEWS
CONTACTS
What I Check When the Fraud Filter Can't Tell a New Business From a Fake One | Siva Balasubramanian
Personal Anecdote 07

The fraud model caught more fakes every quarter.
It also caught more real businesses.

07
Marketplace Trust & Integrity · Personal Anecdote

What I check
when the fraud filter can't tell a new business from a fake one.

catch rate climbing for three quarters, real-business appeals climbing right alongside it →

Our fraud model got better at catching fake business listings for three straight quarters. It also got better at suspending real, newly-opened small businesses that happened to look like fraud on paper. I almost approved a fourth tightening pass before I built the habit of checking who a "win" was actually landing on.

Almost missed itTook 3 tries to get the tiers rightNow a habit I default to
Personal Anecdote

This is a Personal Anecdote, not an employer engagement. Written to show how I diagnose and structure a trust & safety product problem in a two-sided local marketplace, the numbers below are original.

0
Enforcement tiers designed
0
People on the team
0
Weeks to ship
0
% cut in false suspensions
Ch. 01 · The Check

What I Check

A fraud model can be winning and quietly punishing the wrong people at the same time

I almost approved a fourth threshold tightening without a second look. The dashboard was green, the trust & safety team was proud of the catch-rate gains, and I had two other launches competing for my attention that week. What stopped me was a queue nobody on the review was required to open.

Aggregate precision on a fraud model is dominated by whatever population is easiest to classify, the obvious fraud rings, so it can improve release after release while a smaller, real population of legitimate small businesses quietly gets misclassified underneath it, and nobody's dashboard is built to catch it. So now, before I sign off on any trust & safety threshold change, I check the false-positive rate for the segment least equipped to appeal, separate from the aggregate catch rate, every single time. That one habit is the whole fix. The rest of this case is just the story of the one time I almost skipped it.

Here's the enforcement pass that taught me to check, and the four things I built so I'd never have to remember to. 👀

Ch. 02 · The Case

The Case

Role & scope

PM owner for local business listing trust & safety enforcement. Accountable for the enforcement-policy decision, not just the investigation, working across the fraud ML team, trust & safety ops, support, and legal/policy.

Stakeholders aligned

Trust & safety leads (co-designed the enforcement tiers), fraud ML team (owned the risk model), support ops (absorbed the appeals volume), and legal/policy (signed off on the verification path).

Business problem

Three consecutive fraud-model tightening passes looked like wins on the topline (fraud catch rate up an estimated 38% cumulative), but appeals-reversal rate for suspended listings rose from 9% to 24% over the same period, concentrated almost entirely among newly-registered, low-digital-footprint small businesses.

Goals

Prove the false-suspension pattern was real, find exactly who it was hitting and why, and ship an enforcement redesign that didn't reopen the door to real fraud rings or slow response time, without fragmenting a review pipeline that already worked.

The Tension

The same model that got better at catching fake storefronts also got better at looking like one to real ones.

How do you build one filter that keeps out fraud rings actively trying to look legitimate, without also catching the real businesses that look inexperienced because they are?

Every fraud-model tightening pass was, individually, a win, catch rate climbed for three consecutive quarters against a growing wave of fake listings. I caught the problem in a queue nobody on the enforcement review was required to check, where the reversal rate on appealed suspensions had quietly climbed from 9% to 24% over the same period.

I want to be clear I wasn't uncovering some team's mistake, nobody was doing anything wrong by the metric they were optimizing for. I decided the metric itself was incomplete, not the model. Aggregate precision is dominated by the much larger population of obvious, easy-to-classify fraud rings, so a model can improve on average while getting meaningfully worse for a smaller, real population of new business owners whose only crime was looking unestablished, thin review history, a shared strip-mall address, a name that didn't yet show up anywhere else online. That distinction, between a bad model and an incomplete objective, is what determined everything I did next.

The forces bearing down on one enforcement decision

ENFORCEMENT DECISION CATCH RATE LEGIT OWNERS SUPPORT LOAD POLICY RISK

Four forces pressing inward on the same enforcement decision, each meter is how hard that force was pushing, the fix had to hold under all of them at once.

From Undifferentiated Risk to a Named Segment · how the appeals pattern first looked

Undifferentiated risk · week 1

? ? ?
?data artifact in the appeals log
?support just backlogged
?real false-suspension problem

Named segment · week 3

Confirmed: Not a data artifact, not a support backlog
Segment: New, low-footprint legitimate owners
Reversal rate climbed 9% → 24% over three quarters

A hunch isn't a case yet. Time to go find out if it holds up. 🔎

Ch. 03 · The Investigation

Research

What the data actually showed

A hunch isn't a decision. Before I brought this to anyone, I ran four separate cuts of the data to rule out the boring explanations first, a data artifact, a tracking bug, seasonality, before treating it as a real product problem worth a team's time.

[1] AGGREGATE METRICS

"Fraud catch rate up an estimated 38% cumulative, both listing removals and repeat-offender re-registration blocks improved release-over-release across three tightening passes."

[2] APPEALS METRICS

"Reversal rate on appealed suspensions rose from 9% to 24% of all suspensions over the same period, nearly 1 in 4 suspensions was later found to be wrong."

[3] SEGMENT BEHAVIOR

"79% of wrongly-suspended listings shared three traits: registered under 90 days, fewer than 5 reviews, and a business address shared with other listings, a strip mall, a shared kitchen, a market stall."

[4] MANUAL REVIEW

"Of the fraud rings I audited by hand, 61% had deliberately built up an aged account, real reviews, and a unique address, exactly the signals a young legitimate business hasn't had time to accumulate yet."

Once the appeals-reversal signal held up, I ruled out the two most likely operational causes before accepting the harder answer, the risk model itself.

Ruling out · week 2

? ? ?
?reviewers rushing appeals
?one bad model version
?risk score itself is the problem

Confirmed · week 3

!
✕ Ruled out: Reviewer rushing
✕ Ruled out: A single bad model version
Confirmed: Risk score keys on thin footprint, which fraud rings and new legit owners share alike

Illustrative, indexed for shape not scale

Two dashboards, same three quarters 👀

High Low Q1 Q2 Q3
Fraud catch rate
Appeals reversal rate

Both metrics climbing looks fine until you plot the gap between them, that gap is the whole story.

100 reversed suspensions, sorted by business type

Small slice, outsized harm

New, low-footprint (79)
Established (21)

New listings are a minority of overall volume but carry most of the reversed, wrongful suspensions.

Problem statement

The fraud risk score optimizes for aggregate precision in a way that is blind to a real, low-footprint legitimate-owner segment worth an estimated 6% of new listings.

North-star metric

False-suspension rate (target: back to 9% baseline), tracked with equal weight alongside fraud catch rate, segmented by account age.

Guardrail metrics

Aggregate fraud catch rate (must not regress below current 38% cumulative gain); median appeal resolution time (must hold under 72 hours).

Non-goals

Rebuilding the core fraud risk model architecture, this was an enforcement-policy problem, not an algorithm problem.

This is the exact scorecard I brought into the policy review, so the trade-off would be a group decision, not a private judgment call.

2 CATCH RATE 1 FALSE-SUSP. RATE 3 APPEAL SLA
  • North star: false-suspension rate, new low-footprint owners
  • Ring 2 , primary guardrail: aggregate fraud catch rate (must hold flat)
  • Ring 3 , secondary guardrail: appeal resolution time

KPIs I actually tracked week to week

False-suspension rate

↓41%improving

Aggregate fraud catch rate

→+38%flat (target)

Appeals reversal rate

↓16%improving

Median appeal resolution

→61hrsflat (guardrail)

Support ticket volume

→+1wkslower triage (accepted)

Verified fast-track adoption

↑9ptsimproving

Four ways to fix it. Only one of them didn't require reopening the door to fraud.

Ch. 04 · The Decision

Ideation & Decision

Weighing the alternatives

Four ways to close a 15-point false-suspension gap, each with a different cost. I scored them on the same two axes so the trade-off would be legible to people who hadn't spent three weeks in the appeals data with me.

1

Keep tightening the same binary risk threshold. $0 cost, but provably blind to a false-suspension harm already at a 24% appeals-reversal rate.

2

Full manual review of every flagged listing. Most accurate, but needs an estimated 40 additional reviewers, doesn't scale with listing growth.

3

Move to tiered enforcement plus a legitimacy fast-track.Chosen ~9 weeks, 0 new headcount. No new model, just a change to how the same risk score gets acted on.

4

Broadly loosen the risk threshold. Cheap, but lets more real fraud rings through, unacceptable against the platform's core trust promise.

Accepted riskTiered enforcement adds roughly a week of triage lag on the lowest-severity tier. Accepted deliberately, a slower soft-warn beats an instant wrongful suspension for a business that can't easily appeal.
Deliberately not builtFull manual review of every flagged listing, would have required roughly 40 new reviewers to keep pace with a queue growing faster than headcount, for a problem tiered enforcement solved without adding a single reviewer.

I brought this scorecard into the room instead of a recommendation memo, so trust & safety could see me rule out option 4 on evidence, not push it aside on opinion.

Weighing options · week 5

? ?
?loosen thresholds: fewer false blocks, but more fraud?
?tiered enforcement: safer, but does it slow us down?

Chosen path · week 6

✕ Loosen thresholds (rejected): lets more fraud through
Tiered enforcement (chosen): catches fraud, protects legit
Same two directions, now impossible to misjudge
My Mental Model

Scalability vs. Trust Preserved

Four options, one question: which protects the most real businesses without opening the door back up to fraud? Full manual review feels like the safe choice because a human looks at everything, but it doesn't scale, and a queue that grows faster than headcount just becomes a slower version of the same harm.

Tiered enforcement is the only option that scores well on every axis at once, without adding a single reviewer.

OptionScalabilityTrust preservedSpeed to ship
1. As-isHighLowInstant
2. Manual reviewLowHighSlow
3. Tiered ChosenHighHigh9 wks
4. LoosenHighLowInstant

Deciding was the easy part. Getting a whole team to actually want the tiers was the real work. 🛠

Ch. 05 · The Build

Solution

What if the enforcement system stopped treating uncertainty as guilt?

Actors & Inputs

  • Business owners
  • Fraud risk model
  • Trust & safety ops
  • Listing signals by account age

System Change

  • Tiered enforcement added
  • Verification fast-track for new owners
  • Explicit appeal SLA

Outcome

  • Real businesses stay live while appealing
  • Fraud catch rate protected
  • Pattern reused on reviews team

From Draft Logic to Confirmed Policy · how I thought through the tier logic itself

Draft logic · week 7

? ?
?if risk high AND account new -> suspend??
?who owns the verification path

Confirmed policy · week 8

W R S
Warn: Risk flag, no verification attempt yet
Restrict: Risk flag, verification offered and ignored
Suspend: Repeat-offender pattern or failed verification
Owned jointly by PM + trust & safety, not a silent gate

Type the business name. Then toggle to see what the owner actually saw.

Mariana's Alterations, opened 41 days ago
Listing suspended. No reason shown, no appeal path visible.
Restricted, verification requested
Appeal open, 5-day SLA

I underestimated one thing going in: the trust & safety team's incentives were built entirely around fraud catch-rate targets used in their own quarterly reviews, so when I proposed a change that could show up as a lower headline catch rate, I watched it land as a criticism of their work, not what I actually meant it as: a gap in what the metric was allowed to see.

Trust & safety wanted

Maximal fraud-catch as the sole metric, worried any softening reopens the door

vs.

PM pushed for

Equal weight on protecting legitimate owners with less power to appeal

Resolved by: a false-suspension metric both teams agreed to co-own
I didn't bring a mandate to soften enforcement; I brought the appeals-reversal data and asked trust & safety to help design the tiers themselves, so the policy was theirs to defend, not mine to impose.How I led

How the quarter actually went

WEEK 1

Signal noticed in the appeals queue

Reversal-rate climb flagged during a routine ops review, unrelated to any model launch.

Noticed
WEEK 3

Root cause traced to the risk score, not enforcement execution

Manual audit confirmed most reversed suspensions were real businesses; the score just couldn't tell them from fraud.

Diagnosed
WEEK 6

Tiers proposed, thresholds co-designed

Trust & safety team helped set the tier boundaries instead of having them imposed.

Proposed
WEEK 9

First warn-tier threshold too loose, recalibrated

A genuinely active fraud ring nearly stayed in warn status too long, threshold tuned.

Adjusted
WEEK 9

Shipped, then extended to a second surface

Tiered enforcement live for listings; reviews-integrity team adopted the same pattern.

Shipped 🎉

Proof

HypothesisA tiered enforcement policy with a verification fast-track will reduce false suspensions of legitimate new businesses without letting more real fraud through.
DesignFlagged listings evaluated against both the aggregate fraud-catch benchmark and the new false-suspension guardrail before an action was taken automatically.
Primary metricFalse-suspension rate for newly-registered, low-footprint listings.
Result False-suspension rate dropped 41% (24% → 14% of flagged listings); aggregate fraud catch rate held at +38% cumulative, not regressed; median appeal resolution slowed by roughly 1 week of triage lag on the lowest tier.
Follow-upSame tiered pattern adopted by the reviews-integrity team the following quarter, protecting an additional ~5% of new listing volume with 0 new headcount.
Tiered enforcement with a verification fast-track
Legitimate new businesses stay live to appeal
Fewer wrongful suspensions, same fraud caught
Marketplace trust holds for both sides at once

Shipped isn't the same as done. Here's what I'd actually keep, and what I'd do differently. 🕑

Ch. 06 · The Retrospective

Learnings

Key takeaways

What worked

I didn't touch the fraud model architecture, I just made a signal we already had impossible to ignore at enforcement time, at roughly zero incremental headcount versus the manual-review alternative.

What didn't

I set the first warn-tier threshold too loose, and it nearly let a genuinely active fraud ring sit in warn status too long in week 9.

Next time

I'd pilot the verification fast-track for a narrower cohort first, so I could calibrate the tier boundaries with less friction and fewer missed fraud rings.

The pattern outlived the project: the reviews-integrity team adopted the same tiered-enforcement pattern the following quarter, protecting an additional ~5% of new listing volume without adding headcount, which is the part of this I'm proudest of, not the fix, the fact that it became a default other teams reached for.

"A catch rate that doesn't count who it wrongly caught isn't measuring trust, it's measuring volume."

Principle carried forward

"Vision without evidence is theater; evidence without vision is local optimization."

Principle in practice

Frameworks & skills applied

North-star & guardrail design Segment-level evaluation Trust & safety policy design Cross-functional alignment A/B experimentation Scalability vs. trust prioritization North-star & guardrail design Segment-level evaluation Trust & safety policy design Cross-functional alignment A/B experimentation Scalability vs. trust prioritization
Ch. 07 · The Fine Print

Quick Answers

FAQ

Is this a real project from a specific employer?

I wrote this as a Personal Anecdote to show how I think through a problem in Marketplace Trust & Integrity, not to claim a specific past engagement.

Why tiered enforcement instead of a better fraud model?

Once I dug in, the model looked reasonable given its constraints, I decided the enforcement policy built on top of it was the actual gap, so that's what I fixed.

What would you do differently?

I'd pilot the verification fast-track on a narrower cohort before rolling it out broadly, I'd have calibrated the tier boundaries with a lot less friction.

How did you get buy-in without formal authority over trust & safety?

I didn't bring a mandate, I brought the appeals-reversal data and asked the trust & safety team to help set the tiers themselves. A policy a team co-designs is one they'll defend; one imposed on them is one they'll route around.

Turning ambiguous product problems into launch-ready decisions.

Say hi →