All work

AI · SaaS · 2022 — 2025

Shipping LLM features on a platform people already pay for.

Craftly.AI is an AI content platform built on custom LLMs. As Technical Project Manager I managed product development across LLM-powered features and third-party integrations — sequencing fast-moving model work against the stability a paid product owes its users, and keeping the technical trade-offs visible to business stakeholders rather than buried in engineering.

Technical Project Manager · Craftly.AI — Codingcops · Visit the live product ↗

LLM
Custom-model features shipped into a live product
Continuous
Release cadence held while the model surface moved
3rd-party
Integrations managed alongside core feature work
Full
Agile ceremonies run across engineering, design and QA
Craftly.AI homepage — 'Changing the way companies write', an AI copywriting platform built on custom LLMs.

Executive Summary

The tension is permanent: the model moves, the product has to hold still.

An AI content platform lives with a contradiction. The model layer changes fast and rewards moving with it. The product layer is something customers pay for monthly and expect to behave the same way tomorrow as it did today.

Managing that is not a one-time architectural decision — it is a sequencing job that recurs every sprint. Which model work is worth the disruption, what has to be insulated, and what ships behind a flag.

I ran product development across LLM features and third-party integrations, kept engineering, design and QA on a single cadence through the full set of Agile ceremonies, and made sure the technical trade-offs were legible to the people making commercial decisions.

At a glance

Role
Technical Project Manager
Product
Craftly.AI
Organisation
Codingcops
Domain
AI content · SaaS
Owned
Product development, delivery cadence
Partners
Engineering, design, QA

Section 02 — The Problem

A moving model surface under a subscription product.

Stability

Paying users notice regressions, not improvements.

A model change that improves average output but alters behaviour a customer had built a workflow around registers as a break, regardless of the benchmark.

Pace

Standing still in AI is its own kind of risk.

The capability floor rises constantly. A platform that insulates itself completely from model progress is stable right up until it is obsolete.

Translation

Business decisions were being made without the trade-offs.

Commercial commitments depend on technical constraints. When those constraints stay inside engineering, the commitments get made without them.

Section 03 — Research & Discovery

Discovery here is continuous, because the ground keeps moving.

On an AI product, discovery is not a phase that closes before build. What the model can reliably do this quarter is different from last quarter, which means the question of what is worth building is permanently open.

Practically, that means capability review has to have a standing slot rather than happening when someone notices. A model improvement is a backlog input like any other — it competes with feature work on the same list, and treating it as an interrupt is how AI products lose their release rhythm.

Section 04 — Current State Analysis

Two kinds of work, sequenced against each other.

Work typeRisk profileSequencing rule
LLM feature workOutput quality can shiftShip behind validation before it becomes default
Third-party integrationsExternal dependency, external timelineTrack as a dependency, not an estimate
Platform / stability workInvisible when it succeedsReserved capacity, not requested each sprint
Prompt / quality tuningImproves one case, can regress anotherNever ships without the earlier cases re-run
Work types and how each was sequenced

Section 05 — Competitive Research

In a category moving this fast, most competitor moves are noise.

AI writing was one of the most crowded categories of the period, with a new entrant most weeks. The discipline is in not treating every launch as a signal — chasing feature parity with everything shipped nearby is how a roadmap dissolves into reaction.

The filter used was whether a competitor move changed a user expectation rather than merely added a capability. Expectation shifts have to be matched, because users arrive already assuming them. Everything else is somebody else's bet, and it stays theirs.

The filter

Match expectation shifts. Ignore capability announcements.

When a competitor ships something that changes what users assume any tool in the category can do, that becomes table stakes and has to be matched. When they ship a capability users have not started expecting, it is their bet, and following it is how a roadmap becomes reactive.

The cost of following

Every matched feature is a stability budget spent.

On a paid product, chasing parity is not free — it consumes exactly the capacity that keeps existing behaviour trustworthy. Declining to follow was usually a decision about what the platform owed current customers, not about the feature itself.

Section 06 — Key Insights

Three things that made the cadence sustainable.

Insight 01

Model changes need a rollback story before they ship.

Non-deterministic output means a change cannot be fully proven in advance. What can be guaranteed is the ability to reverse it quickly when a customer says it got worse.

Insight 02

Integrations are dependencies, not estimates.

Third-party work runs on somebody else's timeline. Treating it as a task with a story point is how a sprint quietly fails on something nobody controlled.

Insight 03

Trade-offs stated in business terms get decided properly.

The same constraint framed as latency versus quality is an engineering detail; framed as cost per request versus churn risk it becomes a decision the business can actually make.

Section 07 — Design Strategy

Four rules for shipping AI into a paid product.

  1. Nothing model-facing ships without a way back.

    You cannot fully predict output quality, so guarantee reversibility instead of certainty.

  2. External dependencies are tracked, not estimated.

    A third-party timeline is a risk to manage, not a number to put in a sprint.

  3. Translate every trade-off into business language.

    Constraints that stay inside engineering get overridden by commitments made without them.

  4. Protect a stability budget every cycle.

    Feature pressure is constant and platform work is invisible until it fails, so it needs reserving rather than requesting.

Section 08 — The System

Every model-facing change has an exposure level.

Internal

Behind a flag, visible to the team only. Output is being assessed before anyone outside sees it.

Live

Default behaviour for customers, with the previous behaviour still reversible.

Deprecated

The old path is retired. On a subscription product this step needs notice, because someone has built a workflow on the behaviour being removed.

Section 09 — The Features

What actually shipped.

Craftly.AI homepage showing the product interface — a writing assistant with a rephrase control and a project dashboard for generated content.
Product

LLM-powered content features on a live SaaS platform.

Product development managed across custom-model capability and the platform around it, delivered on a continuous cadence with engineering, design and QA on one rhythm.

EditorProjectsBrand voiceAPIPublishing toolsAuth providersBilling
Integration map — the platform's own surfaces on one side and third-party publishing, auth and billing services on the other, every connection crossing one API layer tracked as a dependency.
Integrations

Third-party integrations delivered alongside core features.

External connections managed as tracked dependencies with their own risk profile rather than folded into feature estimates.

Section 10 — User Flow

The product is a loop, and the exit points are where the value is.

  1. The user describes what they need written.

    The first exit point. If framing the request is harder than writing the thing, the product has already lost.

  2. The model generates a draft.

    Where a custom LLM earns its place — the output has to be close enough to be worth editing rather than restarting.

  3. The user judges it and either keeps or refines.

    The second exit point, and the one that decides retention. A rephrase is a success; a rewrite from scratch is a failure the metrics may record as usage.

  4. The draft leaves for wherever the work actually happens.

    The loop only counts as complete when something leaves it. A tool that produces drafts nobody exports has generated text, not value.

Which of the two exit points loses more users is the single most valuable number on a writing tool. Losing people at framing means the input model is wrong; losing them at judgement means the output is not close enough to edit — and those two failures need completely different fixes.

Section 11 — End-User Experience

People do not want AI output. They want to stop staring at a blank page.

The job is not producing text — a model does that trivially. The job is getting someone from nothing to a draft they are willing to edit. Output that is technically fine but not worth editing has failed at exactly the thing the product exists for.

The measure that matters is time to a draft worth editing, not volume of text generated. Those two metrics can move in opposite directions, and optimising the wrong one produces a product that looks busy and gets cancelled.

Section 12 — Impact & Outcomes

LLM capability shipped continuously into a live product.

LLM
Custom-model features delivered into production
Continuous
Release cadence sustained across the build
3rd-party
Integrations delivered alongside core work

Stated honestly, the outcome here is a sustained delivery cadence on a product where the underlying capability kept moving — and cross-functional delivery that held together across the full build.

Release and quality figures are not published, so none are claimed. The one worth tracking on a product like this is rollback rate on model-facing changes: it is the honest test of whether the reversibility discipline was real or just stated.

Section 13 — Reflection

What I would do differently, and what is still open.

Would do differently

Define output quality before shipping against it.

Model-facing work was managed with reversibility as the safety net, which is sound. An agreed quality bar — even a rough one — would have turned some of those judgement calls into decisions with an answer.

Still open

Shipping steadily is not the same as shipping the right things.

A sustained cadence proves the delivery system worked. Whether the features chosen were the ones that moved retention is a separate question, and it needs product data rather than delivery data to answer.

WhatsApp