
You open Claude Code on a Friday evening with an idea for a product. By Sunday there is a login screen, a dashboard and a database that saves real records. Three weeks later the project is still open in a terminal tab, and every new session seems to break something that worked the week before.
This pattern is common enough to have a recognizable shape. AI coding agents such as Anthropic’s Claude Code and OpenAI’s Codex are good at producing the first version of a product. The work that comes after it, turning a demo into something customers can rely on, is where many projects slow down and stop.
Developers feel this too. In the Stack Overflow 2025 Developer Survey, 84% of respondents said they use or plan to use AI tools. Their most common frustration, named by 66%, was “AI solutions that are almost right, but not quite,” and 45% said debugging AI-generated code takes more time.
This article looks at why projects started with Claude or Codex stall, which parts of a product tend to be left unfinished, and how to get a half-built project to launch, with or without outside help.
Why do projects started with Claude or Codex stall halfway?
A project usually stalls when each change starts costing more than the progress it buys. In the first days the agent works in an empty folder, so every request adds something you can see. A few weeks in, every request lands on code that already exists, and each change has more ways to break something else.
Speed can also be harder to judge than it looks. A 2025 randomized study by METR found that experienced open-source developers took 19% longer to finish tasks when they were allowed to use AI tools (Cursor with Claude 3.5 and 3.7 Sonnet), while believing the tools had made them 20% faster. The study measured early-2025 tools on large, mature codebases, and models have improved since. The gap between how fast work feels and how fast it moves is still worth keeping in mind as a project grows.
The agent forgets what was decided last week
Claude Code and Codex work inside a context window, which holds everything the agent knows about your project during a session. Anthropic’s Claude Code guidance notes that performance degrades as the context fills, and that the agent may start forgetting earlier instructions. When a new session starts, earlier decisions are gone unless someone wrote them into the repository.
The result is drift. One session stores dates as text and the next stores them as timestamps. One screen uses the shared button component and another gets a new one. Each choice is reasonable on its own, and together they make the codebase harder to change.
Code that is almost right piles up
Code that is almost right passes a quick click-through. The gap shows up later as an edge case: an empty list, a slow network, a user who signs up twice with the same email.
Without tests, the agent has no way to check its own work, so a fix for one edge case can quietly break another. Anthropic’s guidance puts it plainly: “Claude stops when the work looks done.”
Nobody wrote down what the product must do
Many AI-started projects begin with one prompt and grow one feature at a time. There is no written description of the users, the core journeys or what “done” means for the first release. The agent fills those gaps with reasonable guesses, and the team finds out which guesses were wrong only when they test the app.
Which parts of an AI-built project are usually unfinished?
In the AI-built projects we are asked to take over, the screens people can see are usually in good shape. Most of the unfinished work sits underneath them.

Authentication and permissions. Sign-up and login work, but roles, password resets, session expiry and the rules for what each user can see are often missing.
Data model and migrations. The database schema changed many times during prototyping, and there is no safe way to change it again once real users have data in it.
Payments and integrations. Webhooks, retries and failed-payment states rarely appear in a demo, yet they decide whether your records and your payment provider stay in sync.
Error handling and empty states. The main path works. Slow connections, failed requests and new users with no data often see blank screens or raw error messages.
Security. Veracode’s 2025 GenAI Code Security Report found that 45% of AI-generated code samples failed security tests and introduced OWASP Top 10 vulnerabilities. Secrets in client code, missing input validation and open database rules come up often.
Deployment and monitoring. The app runs on a laptop. Production needs separate environments, backups, logs and a way to know when something breaks.
Tests. Without automated tests, every change has to be checked by hand, and the agent cannot verify its own work.
None of these are unusual problems. They are the ordinary work that separates a prototype from a product, and they are hard to describe in a single prompt.
Should you start over or finish what you have?
When a project stalls, opening a fresh session in a new folder is tempting. An empty project feels fast again, the same way the first one did.
A restart rarely removes the cause. If the requirements, the decisions and the checks were missing the first time, the second version tends to stall at the same point. A rewrite also throws away the parts that already work, which in most AI-started projects is a large share of the code.
A better first step is to find out what you have. Go through the codebase and sort each part into one of three groups: it works and can stay, it works but needs fixing before launch, or it has to be replaced. In our experience the last group is usually the shortest.

How do you finish a project you started with an AI coding agent?
The steps below work whether you keep using Claude Code or Codex yourself or bring in a team. Each one makes the next easier for both people and agents.
Write down what the first release must do
List the users, the three to five journeys they must be able to complete, and what is out of scope for launch. This document becomes the reference for every session that follows. For larger features, Anthropic suggests asking Claude to interview you and write a complete spec to a file before implementation starts.
Give the agent a memory it can read
Decisions that live only in a chat history disappear when the session ends. Claude Code reads a CLAUDE.md file at the start of every session, and Codex reads AGENTS.md files before it starts any work. Use them for the rules that apply to every change: how to run the app, which patterns to follow, which folders to leave alone and how to test. Keep them short, because long instruction files make it easier for the agent to miss the rules that matter.
Add tests around the journeys that already work
Before changing anything else, write tests for the core journeys that work today. Tests protect what you already have, and they give the agent a pass or fail signal it can run without you. Anthropic describes a check the agent can run as “the difference between a session you watch and one you walk away from.”
Work in small changes you can review
Ask for one change at a time, read the diff, run the tests and commit. Start a fresh session between unrelated tasks so the context is not crowded with old attempts. If the agent fails at the same fix twice, stop correcting it in that session and rewrite the request with what you have learned.
Close the production gaps in order of risk
Work through the unfinished list from earlier, starting with the items that put user data or payments at risk: security, authentication, data migrations and payment handling. Visual polish and new features can wait until the product is safe to put in front of real users.
Get a second pair of eyes before launch
Some problems are hard to see from inside the project. A developer who did not write the code, or a fresh agent session asked to review the diff, will catch issues the original session explains away. If the project handles revenue or customer data, or has a launch date, a review by an experienced team usually costs less than a second rewrite.
Conclusion
Claude Code and Codex have made the first version of a product far cheaper to build. The work that comes after it is still there: deciding what the product must do, keeping decisions consistent across sessions, and testing, securing and deploying what was built.
A stalled project is often closer to launch than it feels. The screens and flows that already work are real progress. What is usually missing is a clear list of what is left, a way to check each change and the production work that prototypes skip.
Vizio.ai helps teams take over products that were started with AI coding tools or app builders. We go through the code, decide what can stay and what needs fixing, and take the product to production without rebuilding the parts that already work.
Have a half-built product? Tell us where it stands today.










