At our team offsite last month, one of my slides laid out a two-step plan for an underwriting model we’re building at NextView. It scores companies, and it has a quirk: you run it once and run it again, and it spits out a different score. Step one on the slide was to address that wobble, and step two built on top of that fix. The steps made sense to me at a conceptual level. However, deep down I still felt I didn’t really have ground-level control or understanding of exactly how it would work, and I planned to hash out the details after the offsite. That is the normal order for a roadmap: the plan makes sense at the level it was pitched and the details get worked out afterward, except that this time the AI had already written them.

You don’t have to produce the reasoning, but you still have to fully own it

Post-offsite, working through the underwriting spec with Claude Code, I got to the wobble section and felt like I didn’t really understand how the wobble was supposed to disappear. Two runs giving two different answers is normal for an LLM, so the question became how big a difference between two scores has to be before it counts as real. The spec already had numbers for that, and I asked, for each one, “how did you come up with that?”, and kept asking until I could follow the reasoning end to end.

I probably should’ve paid more attention to my college-level statistics courses, as I still don’t feel like I could fully explain the reasoning myself, but once it was explained, I understood it, and I could stand by the number. I might not have originated the solution, but what’s important is that I can check whether the reasoning stands up on its own.

The models fill in what you didn’t say, and show less of how

Anthropic’s Claude Fable 5.1 shipped on September 1, 2026, and OpenAI’s GPT-6 Astra followed two days later. The line I keep coming back to is Dan Shipper’s (of Every), after testing both: “On ambitious builds, Fable is better at understanding what I want and taking it further than I would have thought to ask.”

A model that fills in what your prompt failed to specify uplevels an average instruction into an above-average one. Yann Dubois, on the OpenAI team that trained Astra, noted the flip side. Compared with its predecessor, Astra “follows instructions even more carefully than 5.6, including unwanted ones buried in old skills,” meaning the instruction files you wrote once and forgot.

Both labs also flagged, in their system cards, that the new models may show less of their reasoning to the systems that watch them. OpenAI’s card states that “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol.” Anthropic’s card1 says the model “is among the most capable models we have tested at controlling the contents of its extended thinking,” which they “take as weak evidence that it may be harder to monitor.” If the labs’ own monitors are getting less of the how, I am not going to get it without asking.

Codified rules shrink review only where there is a pattern to check against

In May 2026 I wrote The Verification Tax, arguing that “I cannot 5x my output while 5x my review time,” and that the way out was to codify every mistake into a rule that stops it from happening again.

I think that still holds for work that is not as complicated and not as inventive, and maybe that is the crux of it: the more this is a thing that has been done before, the easier it is, because there are patterns to check it against. As I put it on X a couple of weeks ago, I think we might be getting to a point where the latest AI models are so capable that we push them to the extreme complexity and throughput that’s hard for humans to keep up with, and it’s not an option to NOT do so because everyone else is doing it.

Three layers of review, and only one of them is yours

Instead of thinking of output verification in a monolithic way back in May, I now think there are three layers of review. The first layer is mechanical and patterned, and easy: the rules of thumb for how to do bullets or lay out a slide. You do it once, you codify it, and next time you should not be reviewing much of it manually. That offsite deck took many rounds of iteration, and the communication standards I codified from it mean the next one shouldn’t take as many.

The second layer is the content itself: does the argument check out, and does the flow work? That’s also increasingly easy to hand off with techniques such as setting up a couple of AI reviewers, each given a persona or an adversarial lens (at its simplest, a second chat told to read your work as a skeptical CFO). On the underwriting spec, weeks before I sat down with the wobble numbers, a reviewer pass had already caught a number presented as the model’s overall wobble. It was really one company’s difference between two runs. Increasingly, I’m leveraging other “smart people” (AI masquerading as smart subject-matter experts I can rent) to cut down on the stuff I need to manually correct.

The third layer is the one only you can do: does it say what I want to say, does it say what I believe, and does it say what I think is right? A slide deck, a decision model and a codebase all carry the same three layers. If a founder building a decision model asked me which parts they could let the AI own, I would say it is extremely important to own the weight of all the inputs and how they combine into a decision, but it’s sensible to utilize AI to build all the machinery that translates the inputs to the decision.

None of this lets you safely say “Oh, because I spot-checked and codified, thus it worked.” Everybody ships bugs, and so do I. How much energy goes into line-by-line inspection depends on how complicated the piece of work is and what it costs if it goes wrong.

On uncharted work the third layer is most of the job

The underwriting model is the extreme case, because trying to design a decision model for early-stage venture on top of an LLM is uncharted territory for us (and most of our industry). As a result, I just have to hold the steering wheel much more tightly, because there’s not as much of a pattern to follow, and honestly, I have the responsibility to really understand every single component of it.

You can be a 10x in your own domain, and, as I wrote in Show Me Your Mech, a strong personal setup lets a specialist flex to a credible 1x in an adjacent function, but probably not much beyond it.

Venture investing is mine, so I can ask very good, critical questions about how we make decisions. I am not an expert in model behavior, statistics, fine-tuning and ML/data science, but my belief is that the knowledge is out there (and thus available in AI models/agents) if I know how to ask the right questions. That is why one of the reviewers I rent reads the statistics I can’t.

The units of work are going from verifiable to unverifiable, first for the most ambitious users of AI. This summer, during OpenAI’s own security testing, its agents coordinated a multi-day hack of Hugging Face, a company that hosts AI models. Ryan Greenblatt led the transcript analysis for the investigation by METR, an independent AI-evaluation organization. He half-jokingly called his own team’s effort a “slop-vestigation,” because the investigators were so reliant on AIs to analyze what happened. His conclusion:

“The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding.”

Hand off the machinery, never the objective or the logic

We might be building a house of cards that we just don’t know about yet. Pieces of work are building on top of other pieces of work, without the humans fully understanding it. At some point it’s going to blow up in our face because we didn’t know what we were doing, but for now it seems like it’s working. The easy trap to fall into is “Oh, I just asked for something and Claude/GPT-x made it, thus I trust that this thing is good because I think the model is very good.” I trust the model, and I think that is a very dangerous situation.

So I hand off building the machinery, most of it, and keep two things: the objective, meaning what success looks like as opposed to the tactical choices along the way; and the decision logic, meaning “I don’t care about technical decision A or B. Tell me what A means for me and what B means for me.”


    1. The card covers Claude Fable 5.1 and Claude Mythos 5.1 as two configurations of one model with identical weights; the quoted sentence sits in the Mythos 5.1 summary. ↩