Digest
Here’s what I wrote, listened to, and saved this week, with a few thoughts on why it caught my attention.
Things I wrote
Getting implicit knowledge written down
A lot of what companies need to make agents useful is still in people’s heads. How work gets done, why someone makes a particular decision, which exceptions matter. Writing that down is work we probably should have been doing already.
One way to get started is to have the agent interview you exhaustively. It can build on your answers and take the conversation somewhere a normal process document wouldn’t. Ask it to find your unknowns or challenge your answers. That can surface things you wouldn’t have thought to include on your own.
Getting my head around Recursive Language Models
I’m far from an expert here, but the implementation pattern makes sense to me. Keep a large body of text outside the model’s working context, let it inspect that text through code, and bring back the pieces it needs. Shoveling everything into the context window can make it harder to use the parts that matter.
The terminology is confusing, though. We use “LLM” to mean the model, the application, and sometimes the whole system around it. Some of what gets described as RLM-like behavior already happens inside the tools we’re using.
Alex Zhang’s Language Model Shape explores a related question: what would change if we designed models around the work agents do? He looks at how different input and output formats could suit particular tasks.
The system around a skill matters
Brett Caughran described sharing skills that work well for him but behave differently for someone else. It’s hard to know how much of the result came from the skill and how much came from context or memory accumulated elsewhere.
I think a lot of this comes back to software and process engineering: understanding how the pieces work together and making the result repeatable. But the models are improving quickly, and some of the extra layers we build become unnecessary. You have to be willing to tear things down and rebuild.
Things I listened to
Gabe Stengel on Invest Like the Best
What resonated with me was building things that don’t quite work yet, but probably will as the models improve. Rogo spent time on the domain expertise and context needed to make the product useful when that happened. I think a lot of the moat is in that work.
- Stengel describes early versions that made good demos but weren’t reliable enough to use. Building early could put customers off, even if you were right about where the technology was going.
- He says expanding into public-equity investing requires expertise Rogo still needs to bring in. Having the underlying tools doesn’t mean you understand how a different kind of investor works.
- One example of understanding the workflow: a managing director can email a marked-up deck, as they already would to an analyst. The product fits that habit, and the analyst can see what changed.
How to Build Team Agents on The AI Daily Brief
Nufar Gaspar walks through the decisions involved in moving from an agent one person uses to something a team can work with.
- Shared knowledge and skills, used by separate personal agents, can be enough. People don’t necessarily need one shared agent, especially when their preferences differ.
- Someone needs to own the knowledge. Interview the people who have it, resolve conflicting answers, and decide who approves updates. Otherwise one person’s version can quietly become everyone’s.
- Check whose access the agent uses, who can ask it questions, and where the answer appears. Someone can have permission to retrieve information without everyone in the channel having permission to see it.
More episodes and takeaways are in my Listening Index.
Other things that caught my attention
Posts
- Thariq on reasoning effort: Experiments with what changes at different effort levels. His working approach is to iterate at low or medium effort, then use high effort for verification. More effort can also mean the model makes more decisions on your behalf when the task is underspecified.
- Michael Thiessen’s fuzzy-linter experiment: He breaks coding guidelines into narrow rules that Jev can check after edits, with uncertain findings sent back for another look. At the time of the post, he’d tested the checks but hadn’t yet tested how an agent responds to them in practice.
Articles
- Jev versus frontier models on customs paperwork: Amari tested transport-mode classification on 1,222 shipments. Jev handled the confident decisions, with Gemini handling the rest. On their test, that cost about one-thirteenth as much as Gemini alone, with nearly the same accuracy. OCR quality and trimming long inputs also affected the results.
- Lance Martin on eval design and hillclimbing: A worked example of improving an application one change at a time and inspecting how it’s graded. In one case, the grader expected three errors when the task only asked for one. Some apparent application failures were problems with the evaluation itself.
- Cloudflare’s new CLI: JSON output by default, commands generated from the API schema, and a search command so agents can find an operation without loading the whole interface.
- Sebastian Raschka on text classification and Jev: A detailed walkthrough from older classification methods to current models, with experiments and an explanation of calibration. He separates the ease of copying an API from the difficulty of making a model work across different tasks.
One month without AI
A developer’s account of stepping away from AI coding after he found himself generating changes much faster than he could review them. The work kept piling up.
The experience he describes feels familiar. There’s a dopamine hit from having agents working on a bunch of things. You can stop paying attention to the work and still feel like you’re making progress, even when some of that progress is artificial. For me, it’s about breaking up the work and figuring out where I need to stay involved. I want the benefits of these tools while still feeling like I have control over what’s happening. Finding that balance is important.