Paying users notice regressions, not improvements.
A model change that improves average output but alters behaviour a customer had built a workflow around registers as a break, regardless of the benchmark.
AI · SaaS · 2022 — 2025
Craftly.AI is an AI content platform built on custom LLMs. As Technical Project Manager I managed product development across LLM-powered features and third-party integrations — sequencing fast-moving model work against the stability a paid product owes its users, and keeping the technical trade-offs visible to business stakeholders rather than buried in engineering.

Executive Summary
An AI content platform lives with a contradiction. The model layer changes fast and rewards moving with it. The product layer is something customers pay for monthly and expect to behave the same way tomorrow as it did today.
Managing that is not a one-time architectural decision — it is a sequencing job that recurs every sprint. Which model work is worth the disruption, what has to be insulated, and what ships behind a flag.
I ran product development across LLM features and third-party integrations, kept engineering, design and QA on a single cadence through the full set of Agile ceremonies, and made sure the technical trade-offs were legible to the people making commercial decisions.
Section 02 — The Problem
A model change that improves average output but alters behaviour a customer had built a workflow around registers as a break, regardless of the benchmark.
The capability floor rises constantly. A platform that insulates itself completely from model progress is stable right up until it is obsolete.
Commercial commitments depend on technical constraints. When those constraints stay inside engineering, the commitments get made without them.
Section 03 — Research & Discovery
On an AI product, discovery is not a phase that closes before build. What the model can reliably do this quarter is different from last quarter, which means the question of what is worth building is permanently open.
Practically, that means capability review has to have a standing slot rather than happening when someone notices. A model improvement is a backlog input like any other — it competes with feature work on the same list, and treating it as an interrupt is how AI products lose their release rhythm.
Section 04 — Current State Analysis
Section 05 — Competitive Research
AI writing was one of the most crowded categories of the period, with a new entrant most weeks. The discipline is in not treating every launch as a signal — chasing feature parity with everything shipped nearby is how a roadmap dissolves into reaction.
The filter used was whether a competitor move changed a user expectation rather than merely added a capability. Expectation shifts have to be matched, because users arrive already assuming them. Everything else is somebody else's bet, and it stays theirs.
When a competitor ships something that changes what users assume any tool in the category can do, that becomes table stakes and has to be matched. When they ship a capability users have not started expecting, it is their bet, and following it is how a roadmap becomes reactive.
On a paid product, chasing parity is not free — it consumes exactly the capacity that keeps existing behaviour trustworthy. Declining to follow was usually a decision about what the platform owed current customers, not about the feature itself.
Section 06 — Key Insights
Non-deterministic output means a change cannot be fully proven in advance. What can be guaranteed is the ability to reverse it quickly when a customer says it got worse.
Third-party work runs on somebody else's timeline. Treating it as a task with a story point is how a sprint quietly fails on something nobody controlled.
The same constraint framed as latency versus quality is an engineering detail; framed as cost per request versus churn risk it becomes a decision the business can actually make.
Section 07 — Design Strategy
You cannot fully predict output quality, so guarantee reversibility instead of certainty.
A third-party timeline is a risk to manage, not a number to put in a sprint.
Constraints that stay inside engineering get overridden by commitments made without them.
Feature pressure is constant and platform work is invisible until it fails, so it needs reserving rather than requesting.
Section 08 — The System
Behind a flag, visible to the team only. Output is being assessed before anyone outside sees it.
Default behaviour for customers, with the previous behaviour still reversible.
The old path is retired. On a subscription product this step needs notice, because someone has built a workflow on the behaviour being removed.
Section 09 — The Features

Product development managed across custom-model capability and the platform around it, delivered on a continuous cadence with engineering, design and QA on one rhythm.
External connections managed as tracked dependencies with their own risk profile rather than folded into feature estimates.
Section 10 — User Flow
The first exit point. If framing the request is harder than writing the thing, the product has already lost.
Where a custom LLM earns its place — the output has to be close enough to be worth editing rather than restarting.
The second exit point, and the one that decides retention. A rephrase is a success; a rewrite from scratch is a failure the metrics may record as usage.
The loop only counts as complete when something leaves it. A tool that produces drafts nobody exports has generated text, not value.
Which of the two exit points loses more users is the single most valuable number on a writing tool. Losing people at framing means the input model is wrong; losing them at judgement means the output is not close enough to edit — and those two failures need completely different fixes.
Section 11 — End-User Experience
The job is not producing text — a model does that trivially. The job is getting someone from nothing to a draft they are willing to edit. Output that is technically fine but not worth editing has failed at exactly the thing the product exists for.
The measure that matters is time to a draft worth editing, not volume of text generated. Those two metrics can move in opposite directions, and optimising the wrong one produces a product that looks busy and gets cancelled.
Section 12 — Impact & Outcomes
Stated honestly, the outcome here is a sustained delivery cadence on a product where the underlying capability kept moving — and cross-functional delivery that held together across the full build.
Release and quality figures are not published, so none are claimed. The one worth tracking on a product like this is rollback rate on model-facing changes: it is the honest test of whether the reversibility discipline was real or just stated.
Section 13 — Reflection
Model-facing work was managed with reversibility as the safety net, which is sound. An agreed quality bar — even a rough one — would have turned some of those judgement calls into decisions with an answer.
A sustained cadence proves the delivery system worked. Whether the features chosen were the ones that moved retention is a separate question, and it needs product data rather than delivery data to answer.