Nikhil Ashok

Product case study · Position2

Building an AI platform that made SEO 6–15× faster.

SEO Studio is an internal platform of 11 AI tools — research, audits, content builds and monitoring — that took a 7-person agency team from drowning in assembly work to shipping strategy. This is how a hackathon prototype became the team's daily infrastructure, and what it cost to learn.

A 0→1 internal product case study · Nikhil Ashok · Product & build lead

faster across the three highest-volume deliverables — measured on our own timesheets AI modules across Research, Optimize, Build and Monitor — one shell, one registry person team using it as daily infrastructure — plus external companies on free access

If you're screening for PM skills, this is where they are

If you only have three minutes

The whole case, compressed

The full study is a ~25-minute read built to be interrogated. This is the skeleton — every claim below is unpacked in the acts that follow.

The wedge. Generative search changed what SEO clients need and how fast they need it, while agency deliverables stayed labor-priced. The reframe that drove everything: the bottleneck was never intelligence — it was assembly. SERP pulls, competitor scraping, metric exports and first-draft synthesis ate roughly 80% of the hours on work that was sold as expertise. So we automated the assembly and defended the judgment.

The eleven, operable

Every module, with the assembly work it removed and the judgment it left alone. Pick any one to load it.

TOOL/01 — keyword research automation Shipped ✓

~30 minutes per page. Hundreds of pages, every client, every month. So I wrote the spec and shipped the automation. The clock on the right is honest — just compressed so you don’t have to sit through it.

The team’s quality bar is built into the output, which is why it survived contact with the team.

runs once as it scrolls into view · hold the button to run it again

30:00MIN / PAGE

30:00 → 02:00 per page · −93%

time figures are from our own timesheets · four of them are timed end to end in Act One →

The four decisions that matter most

01

A platform, not a pile of scripts — one shell, one design system, one run-history layer. The marginal cost of tool #11 was a fraction of tool #1, which is the only reason 11 modules exist.

02

The AI drafts; it doesn't decide — deterministic rule engines find what matters, LLMs write narrative, and every output lands in an editable surface with a human gate before anything ships.

03

Unit cost as a product requirement — one test run burned 800,000 Semrush units. Budgets moved from "be careful" to database-enforced ceilings the application cannot bypass.

04

Prove it internally before selling it — the team is the customer first. External access is free today and used to nurture prospects; the paid model comes only once the tools have earned it.

Honest status

The time savings are measured and I'd stake the case study on them. Adoption depth, quality deltas and revenue impact are not yet instrumented — run history and token spend are the only telemetry live today.

the long version starts below · Act One: The Bind →

Act One · The Bind

AI search broke the unit economics of agency SEO

SEO deliverables are labor-priced. Then generative search changed what a deliverable even is.

Keyword research, content briefs, audits, competitor analysis — sold by the hour, priced by the hour, delivered by the hour. AI Overviews and generative engines changed both sides of that equation at once: what clients need (GEO/AEO readiness, agent readiness, citations inside AI answers) and how fast they need it. A quarterly audit cadence stops making sense when the surface you're optimizing for regenerates weekly.

Meanwhile the work itself hadn't changed shape since 2015. An analyst opened eight tabs, exported four CSVs, reconciled them by hand, and then — in the last twenty minutes — did the thinking the client was actually paying for.

The bind, stated plainly

Three doors, two of them traps. Hire more people — margin death. Do less work — client death. Or change what an hour of expert time produces. Only the third door is a product.

Where to play, and how to win

Before writing a line of code I wrote down the strategic frame, because the tempting version of this project — "let's build an AI SEO tool and sell it" — is a different, much harder company. This is the version we could actually win.

Where to play

Our own delivery workflow first — the seven people whose timesheets I could read, whose output I was accountable for, and who would tell me immediately when something was wrong. Not the open market.

How to win

Depth of domain judgment encoded as deterministic rules, not model access. Anyone can call an LLM. Almost nobody can tell it which twenty things to check on a healthcare landing page and what disqualifies each one.

What must be true

Analysts trust the output enough to ship it, unit cost per run stays well under the labor it replaces, and quality holds under client review. All three are testable — Act Six says how.

Where the hours actually went

The per-tool numbers are in the summary above. This is what they add up to across one full deliverable cycle.

One full deliverable cycle · before → after

3.5 hrs → 32 min

Keyword research, the content brief, content enhancement and the on-page check — the four deliverables that make up one page cycle, timed end to end on real client work.

measured on our own timesheets, not a survey · per-tool numbers in the summary above ↑

Act Two · Discovery

I was the user — and that's a bias worth naming

The best thing about dogfooding is that the backlog writes itself. The worst thing is that it only writes your backlog.

What we did, and what it does and doesn't prove

Discovery here was unusually cheap, because the users sat next to me. Three methods, in order of how much they actually changed the roadmap:

Timesheet analysis. Every deliverable type broken down by logged hours across the team. This is what surfaced the ~80/20 assembly-to-judgment split and killed my assumption that content writing was the bottleneck. It wasn't — the research feeding the writing was.

Shadowing. Sitting with analysts through a full keyword-research and a full audit cycle, timing each step. Tab-counting sounds trivial; it's what revealed how much time went to reconciling exports rather than gathering them.

Interviews with all seven analysts. Open-ended, focused on the last time a deliverable ran late and why. Small n, but it's a census of the actual user base, not a sample of it.

A word on rigor

Seven users is a complete population, not a statistical sample — it tells me what patterns exist and how intensely people feel them, and nothing about whether they generalize beyond this team. That's an acceptable trade for an internal tool and a real limitation the moment we sell it externally, which is exactly the risk flagged in Act Seven.

The bias, named

Being your own user accelerates v1 and blinds you to workflows that aren't yours. The team corrected me on exactly this. The lasting consequence is a product principle, not a patch: every AI output lands in an editable surface. Keyword lists are editable, briefs regenerate per field, nothing is take-it-or-leave-it.

The map that became the product

Discovery produced one artifact that everything else hangs from: every deliverable, broken into steps that are assembly (mechanical, automatable) and steps that are judgment (human-only). This is the intellectual core of the platform, and the reason it didn't turn into a content-spam machine.

Assembly Automate this

SERP & competitor scraping

Third-party metric pulls (volume, difficulty, backlinks)

Data normalization & formatting

Content recommendation & creation

Schema & structured-data markup

Judgment Defend this

Strategy & prioritization calls

YMYL / clinical claims

Brand & client voice

Final approval before anything ships

The "is this actually right?" call

We automated the left column with everything we had. We defended the right column just as hard.

From pain to requirement

Every observed pain became a requirement with an acceptance criterion, so "done" was never a matter of opinion. An abridged view:

Observed pain Requirement Acceptance criterion
Eight tabs and four CSV exports per keyword setSingle-input run producing an enriched, ranked keyword setAnalyst goes seed → export without leaving the tool, under 5 min
AI drafts confidently cite links that don't existInternal-link anchors validated against live page textZero suggested anchors whose text isn't present verbatim on the source page
A three-minute wait feels like a crashStep-level progress that survives a refreshClose the tab mid-run, reopen, run state intact
Analysts won't ship output they can't amendEvery AI surface editable and regenerable per fieldNo terminal state in the product is read-only before export
One careless test burned 800k Semrush unitsPer-run and per-day spend ceilingsA run that would exceed budget is refused, not truncated mid-flight

Getting permission to build it

An internal tool has no budget line and no natural sponsor. The path I took was deliberately public: I entered the company's internal hackathon with the tech audit prototype and won. The win wasn't the point — the sponsorship was. It converted a hunch into an executive mandate and a real budget, which is the difference between a platform and yet another script in someone's downloads folder.

Executive sponsor

Wants: margin defensibility and a story for clients.
Gave: mandate, budget, air cover.
Managed by: demoing working software, never slides.

The SEO team (7)

Wants: fewer tabs, no loss of control over quality.
Gave: requirements, brutal feedback, adoption.
Managed by: shipping their pushback as product principles.

Engineering (22)

Wants: supportable systems on the company platform.
Gave: the migration path to Arena.
Managed by: proving demand before asking for their time.

Act Three · The Bet

What we chose to build — and what we refused to

Three principles governed every scope decision. Each has a real mechanism behind it.

01

A platform, not a pile of scripts

One app shell, a shared design system of 20+ components, a unified run history that survives refreshes, and a single tool registry driving navigation. The payoff compounds: new modules inherit persistence, streaming progress and exports for free, so a new capability ships in days rather than weeks. This is the decision that made 19 tools possible with a team this size.

02

The AI drafts; it doesn't decide

Findings that matter are deterministic — rule engines run ~20-category audits and QA checks, and they return the same answer every time. LLMs write narrative, suggest and summarize. Every AI output lands in an editable surface, and page content passes human approval gates.

The self-check, in detail

The article-brief pipeline runs a second model pass that validates the draft's alignment against the primary keyword and regenerates if it has drifted — AI QA'ing AI, with deterministic rules as the referee.

03

Guardrails are the product in YMYL

For healthcare clients the guardrails are the value. Prohibited-claims rules. Reviews only if real — aggregate ratings are never invented. Doorway-page eligibility checks, so a location page cannot be created unless the location is verified and actually offers the service. A four-gate approval ladder — SEO → Clinical → Content → Client — where any content edit resets every downstream approval.

The flagship trust detail

Internal-link anchor text must exist verbatim on the live source page — the system string-checks it before the link is allowed. No hallucinated anchors, by construction.

The first version, ruthlessly scoped

The MVP was one tool — tech audit — built in a hackathon weekend. That constraint was useful. It forced the question "what is the smallest thing that proves assembly can be automated without losing the analyst's trust?" and the answer was not a suite.

In the MVP

URL → HTML pull → error identification → fix recommendation → export. One deliverable, end to end, with the human edit step present from day one.

Deferred

Run history, caching, multi-user accounts, exports beyond CSV. All real needs — none of them the thing being tested.

The kill criterion

If analysts exported the output and then redid the work by hand, the thesis was wrong. They didn't — they edited and shipped it. That's what justified everything after.

How the backlog got ordered

With one builder and a team that needed everything at once, sequencing was the whole game. Reach was easy to ground — I knew exactly how many analysts ran each deliverable and how often, from the timesheet data. Effort I could estimate honestly because I was building it.

Reconstructed, not retrofitted

I prioritised deliberately but did not run a formal RICE model at the time. The table below reconstructs the scoring from the actual build order and the reasoning I applied, so the framework is honest about being a post-hoc articulation of real decisions — not evidence of a spreadsheet that never existed.

Module Reach Impact Conf. Effort Score
Keyword Research14030.92189
Content Research & Brief8430.8367
Content Enhancement7030.83.548
Schema Generator5610.91.534
SEO & GEO Audit4220.8417
Agent-Readiness Audit
new revenue line
2130.6313
Location Page Builder
parked
1430.482.1
Hub & Spoke Builder
parked
720.480.7

reconstructed RICE · score = (reach × impact × confidence) ÷ effort · reach is monthly runs across the 7-person team, from timesheet data · confidence reflects how well we understood the problem when the call was made — the two parked builders score low on confidence and high on effort, which is exactly why they're still Coming Soon

Saying no on purpose

What we deliberately did not build — and why

No auto-publish. A human gate stands before anything touches a client site. Speed is worthless if it ships a mistake at scale.

No all-in-one mega-tool. Small, sharp tools over one monolith. Each does one job well and composes with the rest.

No premature productisation. Internal leverage first. Charging for something the team hadn't yet stress-tested would have bought a support burden instead of revenue.

Parked on purpose. The Spoke Builder remain Coming Soon —eligibility/complexity bar that wasn't worth clearing yet.

§04Act Four · The Economics

The economics nobody puts in a portfolio

Third-party data and model usage are real money, spent per run. Early on, testing the competitor-research module alone burned 800,000 Semrush units — one careless report design, and the meter just ran. Nothing broke, no client noticed, and that's precisely why it was dangerous: an internal tool with an unmetered API key is a slow leak on the exact margin it was built to protect.

That near-miss made unit cost a first-class product requirement. The governance didn't arrive fully formed — it climbed a ladder, each rung a response to a way the previous one could still fail.

Level 1

No caps

v1 shipped fast and trusted the code to behave. It didn't.

Level 2

Per-run caps

Hard ceilings in application code. A single run can no longer exceed its budget.

Level 3

Daily budgets

Budgets reserved and reconciled against a ledger, not estimated after the fact.

Level 4

Database-enforced

Enforcement lives in the database itself — the application cannot bypass it, even by accident.

cost-governance maturity · the lesson generalizes: if a limit can be bypassed by a bug, it isn't a limit

Different models for different jobs

The platform doesn't have "an AI." It has a provider-agnostic client layer, and each task is routed to the model that tested best for it. That routing is a product decision with a cost and quality consequence, so it's made per task rather than per vendor.

GPT models

Content creation

Drafting, rewriting, summarising, expanding briefs into prose. Chosen on output quality for long-form generation at volume.

Claude models

Decision-making

Classification, scoring, alignment checks, the QA pass that decides whether a draft gets regenerated. Chosen for instruction-following and consistency under rules.

The abstraction

Swappable by task

A single client interface sits in front of OpenAI, Anthropic and Google, so a model can be changed for one task without touching the 11 tools that call it.

Plus: caching tiers, so repeat work is nearly free

7dSERP cache
7dthird-party metric cache
30dLLM response cache

"I treated unit cost as a product requirement, not an ops afterthought."

Act Five · Shipping

Shipping without a launch

There was no launch day. There was a release process wearing the disguise of status tags.

Every tool in the registry carries a live status: Beta, Internal, Internal Testing, or Coming Soon. Those tags aren't decoration — they are the release process. A module earns its way up the ladder by surviving real client work, and it can be demoted just as easily.

Beta Internal Internal testing Coming soon

Modules shipped over time

From one hackathon tool to eleven. Phase markers are approximate.

1
2
5
7
9
11
Hackathon MVP · Research + Audits + GEO / Agent + Build tools Today

Who actually built this

The unusual part of this project isn't the AI — it's who wrote the code. The SEO team built the platform ourselves on Railway, with AI pair-programming doing the implementation lifting. That was a deliberate call: the people with the domain judgment were the only ones who could encode it, and waiting for engineering capacity would have killed the project before it proved anything.

That model has a hard ceiling, and I planned for it from the start. A platform built by its users is fast and fragile. Once demand was proven, the company's 22-person engineering team took over the path to production — migrating the platform onto Arena, the company's analytics platform, where it gets the reliability, scale and support an SEO team moonlighting as developers cannot provide.

Phase 1 · Prove it

SEO team on Railway. Domain experts building for themselves, AI-assisted, optimizing for learning speed over durability. Owned by me end to end.

Phase 2 · Harden it

Engineering, 22 people. Migration to Arena. My job shifts from building to translating — requirements, domain rules, and the reasoning behind every guardrail.

The handoff risk

Tacit knowledge lives in the people who built it. Documentation lagged badly (Act Nine) — which makes this migration the moment that debt comes due.

The roadmap from here

Now

Migration to Arena with the engineering team

Promote Internal Testing modules to Beta on evidence

Instrument adoption and quality — the gap Act Six admits

Next

Decide the paid packaging (Act Seven)

Harden the external-access experience for non-expert users

Knowledge-base / client-context layer, if it earns its complexity

Later

Unpark the two builders, or formally kill them

Horizontal scale once single-machine assumptions are unwound

Client-facing self-serve, if the paid motion validates

Live progress was a trust decision, not a tech flex

The AI pipelines take minutes, not milliseconds. So they show live, step-by-step progress, and that progress survives a page refresh — close the tab, come back, and the run is still there. A three-minute spinner is how you lose a user on day one. Making the wait legible was a product requirement, not an engineering nicety.

Act Six · Measurement

One goal, and an honest account of what we can prove

The objective was never "save time." Saving time is the mechanism. The objective was changing what the team spends its hours on.

Objective

Shift the team's time from operational execution to strategic work.

KR 1

Cut hands-on time for the three highest-volume deliverables by at least 5×. Met — 6–15× measured.

KR 2

All 7 analysts using the platform as their default path for those deliverables. Met — adoption is universal within the team; depth of use is not yet instrumented.

KR 3

Ship the GEO/AEO capabilities clients were starting to ask for, before they became table stakes. Met — GEO Readiness and Agent-Readiness audits are live in testing.

KR 4

Reallocate the recovered hours into strategic work rather than more volume. Unproven — this is the KR that actually matters and the one I can't yet evidence.

The metric I'd hold myself to

North Star: share of team hours spent on strategic work — analysis, recommendations, client conversations — rather than assembly. It's the only number that captures the actual goal, because the failure mode of an efficiency tool is real: you halve the time per deliverable and simply do twice as many deliverables, and nothing about the work has improved.

Input metric

Minutes per deliverable, by type. Measured.

Usage metric

Runs per analyst per week, by module. Partially available via run history.

Quality guardrail

Revision rounds per deliverable. Not instrumented.

Cost guardrail

Unit and token spend per run. Measured.

What's actually instrumented — and what isn't

The honest gap

Today the platform records run history and token / unit spend. That's it. I can tell you exactly what a run cost and who triggered it; I cannot yet tell you whether the output needed fewer revisions, whether clients accepted it faster, or how the recovered hours were actually spent. Building the tool was prioritised over measuring it — a defensible call at the prototype stage and an indefensible one now that it's infrastructure. It's the first thing on the Now column of the roadmap.

The experiment that shaped the architecture

The most consequential test wasn't a feature A/B — it was model selection. Running the same tasks across providers surfaced a split that turned into an architectural principle:

GPT models write better; Claude models judge better. Once we stopped asking one model to do both, output quality stopped being a coin flip.

the finding that produced the task-routed model layer in Act Four

The honest caveat: this was evaluated on our own outputs by the analysts who'd use them — practitioner judgment at working scale, not a blind rubric-scored eval. It's enough to justify an architecture decision that's cheap to reverse. It would not be enough to publish.

Act Seven · Go-to-Market

From internal tool to something worth paying for

The tools left the building before the pricing did. That order was intentional.

SEO Studio was built for one team. It is now also in the hands of external companies, free, and used to nurture prospects — a working demonstration of capability rather than a brochure claiming it. A prospect who runs their own site through an audit and gets something genuinely useful has had a different conversation than one who sat through a credentials deck.

Stage 1 · Internal

The team is the customer. Every module is stress-tested on real client work before anyone outside sees it.

Stage 2 · Free, external ← today

External companies use the tools at no cost. The product does the selling; the agency relationship is what's actually being marketed.

Stage 3 · Paid

Two doors: self-serve for teams who want the tools, or a paid engagement where we run them alongside you.

The signal that says this could work

External users reach out to us when they hit problems

That's a small thing that means a large thing. Nobody troubleshoots a free tool they don't intend to keep using — and the moment a prospect asks us for help interpreting an output, the tool has created exactly the advisory conversation the agency sells. The support request is the lead.

The two doors

The paid model isn't live yet, and the pricing is genuinely undecided. What is decided is the shape: two options, reflecting a real split in how buyers want to consume this.

Self-serve

Pricing TBD

Pay and use the tools. For teams with in-house SEO capability who need leverage, not advice. Cheap to serve, but only viable once the product is usable by someone who wasn't in the room when it was built.

Assisted engagement

Pricing TBD

Pay for the tools and we help you use them. Higher touch, higher value, and the natural upgrade from the support conversations already happening. This is the door most existing signals point at.

Who it's for, and what has to be true

The target is new clients — companies the agency isn't working with yet, where the tools open a door that a cold pitch doesn't. That makes SEO Studio a customer-acquisition asset first and a revenue line second, and it's worth being precise about that, because the two imply very different roadmaps.

What has to be true before charging How we'd know
A non-expert can get value without us in the roomExternal runs that complete and export without a support request
Unit cost per external run is comfortably below any plausible priceAlready measurable — spend telemetry exists today
Free access converts to conversations, not just usagefree users → qualified conversations → clients won
Charging doesn't cannibalise the nurture motionA free tier deliberately preserved for exactly that reason

The risk I'd flag in the room

This product was designed around seven expert users who could route around any rough edge. External users can't. The gap between "works for us" and "works for a stranger" is the single largest piece of unpriced work between here and revenue — and it's a product problem, not a pricing one.

Act Eight · Impact

What it changed

Time is measured. The rest is being instrumented — and labeled honestly until it is.

per-deliverable time savings are in the eleven, operable ↑ · what follows is what the time bought

Adoption

7 of 7 team members use it as daily infrastructure, plus external companies on free access. Runs per week and share of deliverables flowing through the platform.

Quality

Revision rounds, client acceptance, audit findings acted on.

Reach

Outputs touch 10+ client accounts across US B2B and healthcare.

The capability that wasn't sellable before

The biggest impact isn't efficiency — it's the offer sheet. Three things the team could not credibly sell a year ago, now productized:

Agent-readiness audits

Scoring a site 0–100 on how accessible it is to AI agents, including agents-file and structured-access checks.

Per-platform GEO readiness

Readiness scored separately for Google AI Overviews, ChatGPT, Perplexity, Claude, Gemini and Copilot — because each cites differently.

Programmatic page generation

Location × service pages at scale, each passing eligibility and QA gates before it can exist.

The honesty rule

Where a number isn't instrumented, it's labeled — never dressed up as a measurement. The time figures are the ones I'd stake the case study on.

Act Nine · Reflection

What's proven, what's still a bet, and what I'd do differently

Validated

Platform beats point tools — the marginal cost of each new module collapsed.

Humans-on-judgment kept quality while multiplying speed.

The deterministic + LLM split prevented the credibility failures that quietly kill internal AI tools.

Domain experts building their own tooling is viable — for a phase.

Still a bet

Whether free external access reliably converts to clients.

Whether the product survives contact with non-expert users.

Whether the recovered hours actually became strategic work.

Whether the parked builders clear their bar or stay Coming Soon for good.

Do differently

Instrument impact metrics from day one instead of reconstructing them later.

Design for durable, recoverable runs from the start.

Keep documentation level with the build, not several modules behind.

Write the prioritisation down while making it, not afterwards.

The honest version

Durability

Some early modules kept run state in memory. A restart lost in-flight jobs. Later modules fixed this with durable, recoverable runs — durability became a requirement only once real work depended on it, which is exactly the sort of thing you should design for before it bites you.

Scale

Single-machine assumptions crept in that would need real rework to scale horizontally. Cheap to write; expensive to unwind — and now the engineering team's problem as much as mine, which is its own lesson about who pays for your shortcuts.

Documentation

Documentation lagged the build — the README covered one tool while eighteen others shipped. The product moved faster than the story of the product, and the bill for that arrives precisely at handoff.

"The hard part was never the AI. It was keeping human judgment exactly where it earns its keep — and metering everything else."

case study by Nikhil Ashok · internal AI platform, built and shipped for a working SEO team

Two case studies

One shows how I de-risk a bet. This one shows what happens when I ship.

PetPal

Concept → validation. How I de-risk a bet before spending a rupee.

Open →

SEO Studio

Execution → shipped platform. Real users, real cost decisions, real guardrails.

You are here

Connect on LinkedIn ↗ nikashok11@gmail.com →

SEO Studio · a product case study by Nikhil Ashok

Internal AI platform for a working SEO team. Figures are illustrative/directional and labeled as such. Last updated August 2026.

WhatsApp