What I check when a feature works almost too well.
- Home
- portfolio
- Product Management
- What I check when a feature works almost too well.
The answer got better.
The business model broke underneath it.
What I check
when a feature works almost too well.
a feature that worked exactly as designed, and broke monetization by doing so →
Our generative answer feature satisfied user intent so completely that it eliminated the click both ads and publishers depended on to exist. Nobody had shipped anything wrong. That was the confusing part.
This is a Personal Anecdote, not a verified employer engagement. Written to show how I think through a generative AI monetization problem spanning Search, Personalization, and Ads, the numbers below are original.
What I Check
A feature can succeed completely and still break the business underneath it
The first revenue dashboard I pulled after launch looked like a bug. Satisfaction was up, session quality was up, and the number that funded most of the roadmap was quietly falling. I spent a day assuming it was a tracking error before I accepted it wasn't.
When a product's entire monetization model is built on a behavior a feature is designed to eliminate, shipping that feature well is what triggers the crisis, not shipping it poorly. So now, before I ship anything that changes user behavior this much, I check what the surface still needs to monetize on once the feature does its job perfectly. That's the whole practice. Everything else here is the one time I found out the hard way what happens when you skip it.
Here's the feature that taught me to check, and the four things I built to act on it. 👀
The Case
Business problem
The generative answer feature was a genuine user-experience win, but it quietly eliminated the click that both the ads business and the publisher ecosystem were built to monetize.
Goals
Design a new way to monetize a surface with no click, recover as much of the at-risk revenue as possible, and do it without regressing user trust, session satisfaction, or publisher relationships.
The Tension
The generative answer satisfied intent so completely that it eliminated the click both ads and publishers depended on to exist.
How do you monetize a surface whose entire value proposition is not making people click anything?
The generative answer feature was, by every user-facing measure, a genuine win: session satisfaction rose, task completion rose, and users specifically praised not having to click through multiple pages to compare options. It was also, by every monetization measure, a quiet emergency, click-through to both organic results and paid ads on the same query categories fell sharply, because the answer had already given users what they needed before they ever saw a result to click.
I want to be clear this wasn't a bug. The feature wasn't cannibalizing a competitor's product, it was cannibalizing its own platform's monetization engine, on purpose, by being good at its job. That's a much harder problem to fix than a regression, because the thing causing it is the thing you actually want.
Where the query volume actually goes
Two-thirds of generative-answer queries resolve without a click at all, the feature's biggest strength is its monetization engine's biggest problem.
From Thinking to Locked-In · how the alert first looked
Model still thinking · week 1
Locked in · the formal problem statement, week 2
Root cause confirmedA hunch isn't a case yet. Time to go find out how deep this goes. 🔎
Research
What the data actually showed
[1] SATISFACTION DATA
"Session satisfaction and task-completion scores up meaningfully on queries that triggered a generative answer."
[2] MONETIZATION DATA
"Ad click-through and organic click-through both down sharply on the same query set."
[3] FAILED EXPERIMENT
"A traditional ad unit placed directly below the answer performed worse than pre-feature placements and drew complaints."
[4] HIDDEN SIGNAL
"Comparison-style queries showed unusually high downstream purchase intent even without a click."
From Thinking to Locked-In · how I thought through the monetization angle
Model still thinking · week 3
Locked in · the analysis I actually ran
Intent resolvedIllustrative, indexed for shape not scale
Where the clicks went, quarter by quarter 👀
The red band keeps growing every quarter, that's the feature working, and the revenue base shrinking, at the same time.
Same feature, opposite direction
Satisfaction up. Ad revenue down. Same queries.
The two lines crossing is the whole story, the same launch drove both.
Problem statement
The generative answer satisfies intent completely enough to eliminate the click both ads and publishers depend on.
North-star metric
Incremental monetizable value recovered per generative-answer query.
Guardrail metrics
Perceived ad bias / trust score; session satisfaction; publisher referral trend.
Non-goals
Restoring click-through to pre-feature levels, that ship had sailed the moment the feature worked.
North Star, up close
- North star: incremental monetizable value recovered per generative-answer query
- Guardrail: perceived ad bias / trust score (must not tip)
- Guardrail: session satisfaction, publisher referral trend
KPIs I actually tracked week to week
Incremental value per query
Trust / perceived ad bias
Session satisfaction
Publisher referral trend
Callout eligibility precision
Advertiser adoption rate
Four ways to fix it. Only one of them didn't require choosing which team loses. 🧠
Ideation & Decision
Weighing the alternatives
Roll back or throttle the generative answer feature. Fastest fix for monetization, but sacrifices a genuine user-value win and cedes ground competitively, rejected.
Insert traditional ad units directly beneath the answer. Already tested and already failed on both trust and performance, rejected.
Build a disclosed, intent-gated generative commerce surface.Chosen A clearly labeled comparison callout, shown only when commercial intent is detected.
Do nothing and wait for the market to adapt. Lowest effort, but this pattern would only spread to more query categories over time, rejected.
From Thinking to Locked-In · how I actually weighed this
Model still thinking · week 5
Locked in · the spectrum I brought to the room, week 6
ChosenTrust risk vs. revenue capture
Four options, one axis: how much monetization pressure are we willing to apply, and what does it cost in trust? The instinct is to assume more aggressive equals more revenue, but the most aggressive option had already failed in testing.
The chosen option sits at the calibrated middle: enough pressure to capture real commercial intent, not enough to tip the scale on trust.
Deciding was the easy part. Getting Trust and Ads to stop optimizing against each other was the real work. 🛠
Solution
What if the answer stayed pure, and the commercial value moved next to it?
Actors & Inputs
- Users
- Generative answer engine
- Retrieval layer (organic + sponsored)
- Policy review
System Change
- Commercial-intent classifier gates eligibility
- Disclosed callout, separate from answer
- Lightweight sponsored-slot auction
Outcome
- Factual answer stays ad-free
- Commercial intent captured separately
- Publishers get new attribution credit
From Thinking to Locked-In · how I thought through the eligibility logic itself
Model still thinking · week 7
Locked in · the rule that actually shipped
Eligible Not eligibleWatch the answer generate. Then toggle to see how the sponsored comparison actually got woven in.
I underestimated one thing going in: Trust & Policy owned the generative answer's credibility and worried that any adjacency to a sponsored unit would help critics claim the whole answer was compromised. When I first proposed the callout, it read as a request to let Ads back in the door, not what I actually meant, giving genuine commercial intent somewhere honest to go.
Trust & Policy wanted
Zero ad influence anywhere near the generative answer's factual content
Ads wanted
Maximum prominence for the new sponsored callout format
How the quarter actually went
Two quarters, six phases, week by week
Satisfaction up, clicks down, on the same queries
Flagged during a routine cross-team review, not a launch retro.
NoticedBolt-on ad experiment backfired
Traditional ad unit below the answer tested worse than pre-feature placements, forced a rethink of the whole approach.
SetbackShared scorecard proposed, Trust and Ads aligned
Reframed as a shared external threat instead of a feature negotiation.
AlignedStaged rollout by query category
Started with comparison-style queries showing the highest latent commercial intent.
Rolled outShipped, new attribution model stabilized publisher trend
Most at-risk revenue recovered; trust and satisfaction held flat, not regressed.
Shipped 🎉Proof
| Hypothesis | Users will accept a disclosed, contextually relevant sponsored callout after a genuinely helpful answer, as long as it feels additive rather than intrusive. |
|---|---|
| Design | Staged rollout by query category, starting with comparison-style queries showing the highest latent commercial intent. |
| Primary metric | Incremental monetizable value per query. |
| Guardrail | Trust / perceived-bias score, session satisfaction. |
| Result (illustrative) | Most at-risk revenue recovered on affected categories; trust and satisfaction held flat, not regressed; publisher traffic decline stabilized after the new attribution model shipped. |
Shipped isn't the same as done. Here's what I'd actually keep, and what I'd do differently. 🕑
Learnings
Key takeaways
What worked
Measuring trust and revenue on one shared scorecard instead of two competing ones kept both teams honest about trade-offs.
What didn't
The first version of the sponsored callout still tested as mildly intrusive on non-comparison queries.
Next time
Ship the intent-gating logic narrower from day one, and expand it only as trust data justified it.
"A feature that destroys the business model underneath it isn't done, it's half-shipped."
Principle carried forward"When two teams are optimizing against each other, the fastest way through isn't a better negotiation, it's a shared scorecard that makes the trade-off visible to both sides at once."
Principle in practiceFrameworks & skills applied
Quick Answers
FAQ
Is this a real project from a specific employer?
I wrote this as a Personal Anecdote to show how I think through a GenAI monetization problem spanning Search, Personalization, and Ads, not to claim a specific past engagement.
Why not just remove ads from the answer surface entirely?
The free product depends on that revenue. Eliminating monetization from a dominant surface threatens the whole business, not just one team's metric.
Why not maximize ad density to recover revenue faster?
Overloading a trust-sensitive surface risks killing adoption of the generative feature itself. A slower rollout protects the bigger opportunity.
What would you do differently?
Ship the intent-gating logic narrower from day one, and expand it only as trust data justified it, instead of covering more query categories immediately.