\n\n\n\n Your LLM Makes a Terrible Ghostwriter and a Great Sous-Chef - AgntHQ \n

Your LLM Makes a Terrible Ghostwriter and a Great Sous-Chef

📖 5 min read•830 words•Updated Sep 18, 2026

Remember GPT-1? You could get maybe three sentences out of it before it wandered off into word salad. That was the whole demo. A machine that could hold a thought for the length of a tweet, and we all pretended not to be impressed while quietly refreshing the page.

Now it’s 2026 and GPT-6 will hand you tens of thousands of words that hang together, or a working program, without breaking a sweat. The constraint moved. It used to be “can the model produce coherent text.” It is now “can you tell the difference between coherent text and text worth keeping.” Those are very different skills, and the second one is where almost everyone is failing.

The README rule that keeps showing up

One line from a developer write-up this year stuck with me more than any benchmark chart: never let the model write your READMEs, docstrings, or comments. Write those yourself, later. And the author meant it literally.

This tracks with what’s actually happening in practice. LLMs in 2026 will happily generate both code and documentation, and developers keep accepting the code while rewriting the prose. That is not developers being precious about their writing. It is a signal about what these models are good for.

Code has a truth condition. It runs or it doesn’t. Tests pass or they fail. You can hand a model a vague spec, get back something plausible, and the compiler will tell you within seconds whether “plausible” was good enough. Documentation has no compiler. A generated README is confident, well-formatted, correctly punctuated, and describes intent it does not have access to. It tells the reader what the code appears to do. It cannot tell them why you chose this over the obvious alternative, or which part is load-bearing, or where the sharp edge is that cost you two days. That information was never in the prompt, so it was never in the output.

So the model produces something that looks like documentation and functions as noise. Worse than noise, actually, because it occupies the slot where real documentation would go and nobody notices it is missing.

2x is the honest number

The framing I keep coming back to is “2x, not 10x.” That is roughly the realistic ceiling on what this tooling does for serious work, and I think people who report 10x are either doing greenfield CRUD or not counting the time they spend reading the output.

Two-x is genuinely a lot. If someone offered to double your throughput on any other dimension you would take it immediately. The problem is that 2x does not feel like 2x while it is happening, because the work changes shape. You spend less time typing and more time specifying, reviewing, and rejecting. To a lot of people that feels slower even when it isn’t, which is why the discourse oscillates between “this replaces engineers” and “this is useless.” Neither camp is measuring anything.

Plan first, then generate

The single most common failure mode is starting with a vague prompt and hoping the model fills in the gaps. It will fill them. Just not with your intentions.

The workflows that hold up in 2026 front-load the thinking. Brainstorm a detailed specification with the model first. Outline it. Argue with it. Get the shape of the thing settled before a single line of implementation gets generated. Then feed examples, because example-driven prompts consistently beat descriptive ones. Showing the model three instances of the pattern you want works better than three paragraphs describing the pattern you want.

This is a real skill and it is not the skill most people practice. Most people practice asking for things. The useful skill is pinning down what you actually mean, which is uncomfortable, because it turns out you often didn’t know.

The direction nobody talks about

Here is the inversion I find more interesting than any generation use case. Write the thing yourself. Your article, your spec, your migration checklist, as plain markdown. Then hand it to the model and ask it to check the whole thing against itself. Find the contradictions. Find the step you skipped. Find the claim in paragraph four that paragraph nine quietly disagrees with.

That is a much better fit for what these systems do well. Verification over creation. The model doesn’t need your intent to notice you referenced a config flag you never defined. Markdown turns out to be a near-ideal input format for this, which is a funny thing to learn about a text format from 2004.

What to actually take away

Pick your model on how it behaves in your loop, not on leaderboard position. GPT-6 and MiniMax M3 are both credible in 2026 and the gap between them matters less than the gap between a planned prompt and a lazy one.

Let the model write the code. Let it audit your prose. Write the README yourself. It is the part of the repository that carries information no model can reconstruct, and it takes fifteen minutes. Spend them.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top