A couple of minutes before the planned 2:35 start, with the Webex room still filling up, Venky — Anand's contact on the DBS side, who was running the session — had a confession about a photograph.
Prologue00:05The haircut that kept coming out the same
The session needed a high-resolution photo of Anand for the circulation email. He didn't have one, and, as he put it later, “who cares about lack of things these days? Venky, use your agent and make it high-resolution.” That took a few iterations. Somewhere in the iterations, Venky had tried something more adventurous.
Nobody in the room would have called this a research question. It was a small, funny complaint about a gadget. But watch what Anand did with it. He didn't sympathise, and he didn't suggest a better prompt. He asked what had worked and what had failed — and then, within a minute, he was handing Venky a hypothesis. He had been running his own experiments on the thinking level of the image model — low, medium, high, extra high — and had found that the agents couldn't tell the difference at the top, but that fine hair texture was one of the few things that did improve. So perhaps hair is exactly where the model's effort level matters. “This is something that I would love to benchmark,” he said, “and see if for different images, does thinking level…”
“Sure, let me check that. I can send you those photographs.”
Venky — and with that, the first experiment was commissioned before the talk had started
That exchange was the whole talk in miniature. An observation, a guess about why, a way to test it, and a person who agrees to go and look. Everything that followed — ninety minutes of browser tabs, chat messages and a few deliberately outrageous jokes — was a longer version of the same four moves.
At 04:04 Venky decided the room was ready. Mayank opened the session and framed the series: “we are upskilling our people towards AI Native.” Sitaram introduced the speaker as someone who “famously calls himself an LLM psychologist” and who studies “how AI systems behave and fail and can be verified.” Then Anand, at 06:23, said what he'd come to say — the talk's entire abstract.
“I have just one message to communicate frankly, which is: use AI in an empirical way.”
Anand, 06:23
What did that mean? “People will say lots of things—you try it. If it works for you, good. If it doesn't work for you, good, you learned.” He then did something that made the rest of the session an experiment too: he asked the audience for their rules.
Act I · Advice is a hypothesis06:23The room writes the hypotheses
“What are your top prompting tips?” he asked. “If somebody asked you, ‘How can I improve my prompts?’, what advice would you give them?” His own offering, dropped into the chat first, was the old standard: add “think step-by-step.” Then the chat window filled up.
Anand's read of the last one: it “touches upon most closely to the whole notion of empiricism. That is, I believe it when I test it.” The rest of the session was a tour of what that sentence costs — which turned out to be very little.
Notice what Anand did not do. He didn't rank the suggestions, and he said several of them probably made sense. What he did was treat each as a claim waiting for a test. And he had a piece of fresh advice of his own to put on the bench.
Experiment 110:14If you write simply, does the model think simply?
Anand had come across a popular tip on X. Someone had complained that they couldn't follow what the latest models were saying, and Andrew Carr had replied that his own fix was to “only report to me in ASD-STE100 Simplified Technical English.” Another poster suggested making it a permanent instruction. It's sensible. Anand himself, he confessed, had been struggling to understand the models.
“But then I started thinking, hold on, I'm saying write simply; what if it also starts thinking simply? How about we actually test it?” So he did: a handful of real questions, each asked with and without the suffix, then graded by a model on correctness, key drivers, mechanism, caveats, calibration and actionability, with every pair judged in both orders to remove position bias.
Guess before you scrollAcross six tasks, each graded on a rubric of several dimensions, how often did “Answer in ASD-STE100” improve the answer?Click to reveal →
Almost never. In the published grid, the simplified answer lost on nearly every cell — only a few ties and two lone wins in the dozen judged pairs. The simplified version also checked fewer sources on the tasks where that was recorded.
“Against almost every single one of these criteria, adding the prompt ‘Only report to me in ASD-STE100’ made the result worse.”
The fix he derived cost nothing: don't simplify the thinking; simplify the report. “You let it think however it wants, and then, after it finishes thinking, you say, ‘Now explain whatever you thought in simple English.’” Rather than adding it to every prompt or to the custom instructions, he'd paste it in once an iteration finished. “That comes simply from benchmarking.”
The write-up of the experiment, with the full grid of twelve judged comparisons. Open in a new tab → · the test used ChatGPT, six tasks; the author suggests re-testing in a few months.
Experiment 213:29Does it help to praise, shame or bribe the model?
The second family of folklore is more theatrical. Tell the model it's the world's best expert at mental multiplication. Tell it that even your five-year-old could solve this, so stop being lazy. Praise it. Promise it $500. “Do they work? Do they not work? Where do they work? Where do they not work?” Anand's response, again, was that this is “not very hard to benchmark.”
He ran ten prompt styles — reasoning, emotion, politeness, expert persona, incentive, bullying, shaming, fear, praise, curiosity — against 40 models on multiplication problems, ten cases per prompt and model. The tally, from the published page: “Think step by step” was the only variant that helped overall, by about 3.5 percentage points, and only just short of conventional significance. Emotion and shaming slightly hurt — the emotional one made 21 comparisons worse and 7 better. Everything else was indistinguishable from noise. In the talk's words: “interestingly, emotional prompting actually on the margin hurts.” Many of the models, he said, “seem to be panicking.”
The caveat arrived in the same breath, and it is the more important lesson. “This experiment was run before models started shifting to reasoning by default. So if you tell a reasoning model to think step-by-step, I suspect it'll have no impact whatsoever.” The page agrees: the nudge helped the old non-reasoning models and actually hurt some reasoning ones. Even a result that survived the test came with an expiry date.
“Emotion Prompts Don't Help. Reasoning Does” — ten prompt styles, 40 models, 10 cases each. The talk's shorthand counted these as models; the page counts model-by-test-case comparisons. Open in a new tab →
Experiment 315:36Eighty years of mathematicians' advice, audited
Then Anand turned the same instinct on humans. George Pólya had given mathematicians a list of 15 problem-solving heuristics in How to Solve It. “Good advice for mathematicians, and people have been using it for a long time.” Work backwards from the desired conclusion. Enumerate all the cases. Look at the extremes. Had anyone ever measured whether the rules work? The “Pólya Audit” did — 6,747 runs, three models, fifteen heuristics, seven areas of mathematics.
The result was a patchwork. “Working backwards helps with counting-related problems, but it doesn't help with number theory-related problems,” even though you might say counting is very similar to number theory. Case analysis lifted pre-algebra and counting and was “completely useless — in fact, hurts the results — for precalculus.” Number theory resisted nearly every heuristic. Whether the advice helps flips with the domain and the model.
“The whole notion of general advice is something that we can start putting to rest.”
Anand, 17:05
Read that sentence carefully; he did not say advice is worthless. “Things may work in specific situations.” The target is universal advice. The good news is that AI-related advice is unusually testable — “you try it out on a bunch of situations, see if it works, and if it doesn't work, drop the advice.” And agents can generate the candidates too: ask a model for ten ideas, in the direction Vishnu Kiran had suggested in the chat, then do what Venkateswarlu said and test them on a golden dataset. “Benchmark empirically.”
“The 80-Year Blind Spot: The Pólya Audit” — March 2026, 6,747 experimental runs. Open in a new tab →
The hypothesis ledger, after twenty minutes
Advice offered → what happened when someone tested it · results specific to the models and tasks tested
| The advice | The test | What survived |
|---|---|---|
| Think step by step | 40 models, multiplication, 10 prompt styles | Modest, dated Only variant that helped overall (about +3.5 points); mostly on older non-reasoning models. |
| Praise it, bribe it, give it a persona | Same experiment | Nothing to see Expert persona, incentive, politeness: indistinguishable from noise. |
| Tell it how important this is | Same experiment | Slightly harmful Emotion and shaming made more comparisons worse than better. |
| Only report in Simplified Technical English (ASD-STE100) | Six tasks, graded pairs, both orders | Worse thinking Let it think, then ask for a simple explanation. Details → |
| Work backwards (Pólya) | 6,747 runs, 7 maths domains | Depends Helps counting, not number theory. Same advice, opposite effects. |
| Say what not to do; add constraints | Offered by the room, not run in session | Still a hypothesis Exactly the kind of claim to put on your own golden dataset. |
| Keep a golden dataset | The method itself | Anand's pick “I believe it when I test it.” |
Act II · Models19:22The wrong question is “which model?”
“How do you select models?” Anand asked, and the chat answered quickly. Vishnu Kiran: model intelligence — and, later, context window. Sree: complexity, or cost versus time. Jagannath Prasad: domain-specific models. Uttam: compare the output of different models on a common task and choose. Another participant wrote simply, “benchmark and select”. And Siddhartha asked whether the chart Anand was about to show was a “model council” — a discussion among LLMs. (“Not in this case, Siddhartha.”)
Anand's own starting point is a public one: the human-preference leaderboard at Arena, where you ask any question, two anonymous models answer, and you vote. Even the classic joke-question gets a tour: to get to the other side, “but depending on who you ask: philosophers would say this, scientists would say this, Gordon Ramsay would say because it was raw.” He votes for the longer answer. “Most people tend to pick a longer answer just because it's longer,” he admitted, so the benchmark has human biases baked in — but with millions of votes, it's still a reasonable signal.
Then he put that score on one axis and price on the other. The result, in the LLM Pricing chart he keeps updating, is a frontier that moves up and to the left. His metaphor is the part that stuck: in March 2023 the best models were “high school students”; by March 2024, college graduates; March 2025, PhD-level; and by March 2026, “smarter than tenured professors.” (He was explicit that it is storytelling, not psychometrics.) Meanwhile prices fell.
“That's roughly the difference between a $10,000 budget and a $3 million budget for the same thing, and you know which one is more likely to get approved.”
Anand on PhD-level intelligence at about 3.5 cents per million tokens versus about $10, as priced on 1 Oct 2026
Anand's cost-versus-intelligence chart — rankings and prices as of the talk, 1 October 2026. The “high school” to “professor” labels are his metaphor. Open in a new tab → · Arena leaderboard →
But the interesting part of the chart wasn't any point on it. It was how fast the points move. “Barely a week goes by without the frontier changing.” If so, then picking a model is the wrong skill to practice. And so, at 27:01:
“Maybe the bigger lesson is not how do you choose a model, but rather how do you change when the new model comes?”
Anand, 27:01
The question that replaces it: “how do we keep updating our system with newer models with least effort?” Is your system built so that a better model costs you an afternoon, not a quarter? That's a benchmark too — of your architecture. Which brings us to the model everyone was excited about that week.
Case study28:29Hype, converted into a two-dollar question
Jev had just launched, and, in Anand's words, “everybody's going gaga over it.” It's an unusual kind of model: it doesn't generate text; it classifies. You give it an input and a set of allowed answers — high, medium, low; or good, bad, ugly — and it returns a choice with a confidence, in a few hundred milliseconds, for a fraction of a cent. Anand's question was whether to tell his team to switch to it for classification tasks.
He didn't debate the architecture. He had an agent read up on Jev, then took the BANKING77 dataset of banking-support requests, and asked the agent to run Jev against the models his team could use instead. Then he got ahead of the objection every busy engineer in the room was forming.
“Here I am happily saying benchmark, benchmark, benchmark. But the other question that you would have is, ‘Look, Anand, somebody else is doing this for me. I get the result for free. For my task, I have to sit and spend time and effort. I have 250 other things to do. What do you think, I'm jobless?’”
Anand, 31:41 — arguing with the audience's objection before they could make it
His answer: “this entire benchmark, for instance, was created with me spending five minutes.” He'd opened ChatGPT, given it access to his machine through a local MCP connection, and had a short conversation. Here is what he typed — the five steps, verbatim from the shared chat:
- Research: “What is Jev and how is it different from other other frontier LLMs? How does it compare on relevant benchmarks? How is it priced and is it available on OpenRouter? Is it open weights?”
- Cost first: “I'd like to find out how well it classifies the same BANKING77 dataset vis-a-vis good models… What would it cost? How many can I classify to keep the cost down to $2 overall?”
- Pilot: “OK, just run the 77-case six-model pilot as you have recommended. Share and interpret the results… and re-estimate the cost.”
- Publish: “Create a ~/code/llmevals/jev/ directory with all inputs required to re-run this resumably and idempotently. Also add an index.html that explains the findings in simple, clear language…”
- Commit: “Commit and push.”
The chat itself holds a lesson in verification. The agent quoted a full six-model run at $1.72 — then disclosed that a long tool call had timed out and replayed some requests, billing $0.84 of duplicate calls before it noticed and killed them. The clean results file had exactly 77 × 6 unique rows; the account statement was higher. Even the benchmark-building agent's own work was checked against a count.
The agent-built page, as updated on 28 September 2026 to include GPT-6 Luna: ten models, the same 77 requests, scored against gold labels with no LLM judge. Open in a new tab →
The result was not “Jev is bad.” On this pilot, GPT-6 Astra scored highest (68 of 77), Gemini 3.8 Flash was within one case of it at about a fifth of the price, and Jev scored 58 of 77 — at $0.07 per thousand requests. But GPT-6 Luna, newly released, was cheaper still ($0.06) and two cases ahead. Two cases at n=77 prove nothing. They do establish that, for this task, on that date, price alone was no longer a reason to switch — and Anand's conclusion was exactly as modest: “At least today I have no reason to tell the team to switch, or rather I have to tell them to switch to GPT-6 Luna.” He added that the enterprise-licence friction of onboarding a new vendor could be avoided entirely.
He also kept the door open. Jev's latency is a few hundred milliseconds, so he showed a live demo classifying incoming network requests — flagging path traversal, SQL injection and command injection, including attacks aimed at an LLM that a conventional firewall wouldn't read as hostile. “Maybe there are other reasons to use it, but ultimately all of this boils down to knowing your task.” And then the larger point, which the rest of the session kept returning to:
“The task of benchmarking is very delegatable to agents. We are using agents to figure out how good agents are.”
Anand, 34:53
If the cost of a benchmark falls to five minutes and a couple of dollars, the economics of belief change. “When somebody says, ‘Here is universally good advice,’ you create a small little benchmark to tell yourself whether for your task, this is a good idea or not.” That's a benchmark used adversarially — not as a leaderboard to publish, but as a way to say show me.
Callback35:39Back to the hair
Having shown how cheap a benchmark is, Anand circled back to Venky and the photograph — and to a benchmark with no leaderboard at all. He had wondered what the quality setting on the new image model (GPT Image 2.5, with five levels from low to max) really buys. So he asked ChatGPT to generate the same scenes at every level, look at them, and report whether it could tell the difference. “If so, then I will take a look at it.”
The agent drew a face and zoomed in on the forehead wrinkles. At low they were “kind of okay”; at medium, a little more detailed; at high, clearly detailed. But x-high and max looked no better than high — “I can certainly see more freckles, but that was more a random artifact.” It drew a watch, hoping to catch water droplets (nearly perfect even at low), then a barista, where max just barely beat low on the steam and the reflections. Almost everywhere the pattern held: low to medium, yes; medium to high, sometimes. Hair and fibre were the exceptions where the difference showed.
“I will default to medium when I need any kind of decent quality, and if it involves hair, fiber, cloth, that kind of a thing, I will go to high; otherwise, stick to low.”
Anand, 38:45 — the rule the benchmark bought him. “So in short, again, empiricism.”
Notice where that landed: on hair. It is the same texture Venky's bald-head edits were failing to vary. Whether raising the quality level makes the hairstyle change from person to person is still an open question — Venky had promised to send his photographs, and nothing in the session settled it. The published page is careful in the same way: the generations are stochastic, so this is a practical visual probe, not a proof that the top tiers can never be better. It does say where the benefit shows up — make the subject large, then zoom in on skin, hairs and fabric.
“What GPT Image 2.5 Flare quality actually buys you.” Pick a quality level to see the zoomed crops. Open in a new tab →
“But won't public benchmarks leak into training?”
At 39:39, Sree raised the obvious worry in the chat: public benchmarks may not be reliable because of leakage during training. Anand offered three defences, each with a catch. Niche and public: Simon Willison's famous pelican-on-a-bicycle test seems not to have been trained on, with one possible DeepSeek exception, because providers have little reason to care about so specific a drawing. Human preferences, like Arena: hard to game, though Gemini somehow stayed on top for quite some time — perhaps by gaming for what the lay public likes — “but you could say that, look, that is a relevant benchmark in the first place.” Private held-out sets, like Artificial Analysis: hard to game “unless somebody from that company steals the data and joins another company.”
The real problem, he said, is the opposite one: “the benchmarks are getting outdated because of saturation.” When models score 99.5% and 99.9%, you can't tell them apart. “So the leakage seems to be less of a problem than saturation. At least so far, that's what we have seen.” Another quiet argument for building your own: yours won't saturate, if you keep it hard.
Act III · Agents41:59“Please don't say benchmarking generically”
Next question for the room: how do you choose agents? Not models — agents: ChatGPT versus Claude, Codex versus Claude Code. “Please don't say benchmarking generically. If you are benchmarking, I would love to hear what benchmark you're using.” Sree: “based on level of autonomy and control.” Vishnu Kiran: intelligence, context window, and score at low cost. And Venkateswarlu gave the answer most developers inside a large company will recognise: “We don't have access to multiple agents at work.” Anand's reply: “Which is, yeah, probably the easiest way to choose! You've been given something, and that's all you can use.”
That constraint would reappear at the end of the session, in Venky's closing remarks. But first, Anand's own method — starting with an outside benchmark, then replacing it with his own.
The shift he flagged is from price per token to “cost per task — meaning getting the job done.” A cheap, simple model in a simple harness can finish a straightforward task for less than a smarter, cheaper-per-token model that overthinks and redoes the work. On the outside benchmark, he saw Anthropic's Claude Code at the high-cost, high-quality edge, OpenAI in the middle, and some new cheaper entrants worth watching — as of the day. But, he cautioned, that is still models, not entirely agents.
So the first time he needed to choose a harness, he built the simplest benchmark he could think of: one prompt — create a single-page web app that renders a GitHub user profile and activity comprehensively; read the ID from the URL — given to several combinations of coding agent and model. Then he looked at the results.
“Coding Agents — A Comparison”: one small prompt, several agent/model combinations, a manual quality score, the cost in cents and the time in seconds. A red marker means outright failure. The benchmark is from 2025, as Anand noted in the talk. Open in a new tab →
Two of the tools — Grok-4 and GLM-4.6 inside OpenCode — “didn't even give proper output.” Three of the survivors are worth lining up, because you can see the difference without reading a word of code.
The same prompt, three results, each shown zoomed out to fit. Scores are Anand's manual 0–10 ratings from the 2025 benchmark.
Your own work50:01Why invent a benchmark when you already have a history?
That early GitHub-profile test had two properties Anand singled out. The prompt was tiny. And he could judge the output at a glance: “If you show me these three pages and ask me to score, I can immediately say, ‘Aha, this is better, this is worse.’”
“When you are benchmarking, if you can have the agent generate something which you can assess the quality of at a glance… that is probably one of the most powerful and important things that you should be thinking about.”
Anand, 48:40 — even for a back-end task, convert the result into something you can see
Then, with the clock past midway and a promise to leave half an hour for discussion, he skipped ahead to the idea he cared about more. “When creating a benchmark, why should I have to sit and think about a new benchmark?” So he went back to ChatGPT and said, in effect: look at the tasks I've recently been working on. One of them was the very image-quality page he had just shown — it was loading too slowly, and he'd asked a coding agent to fix it. Another was scraping WhatsApp. A third was making an MCP server's console logs readable. The agent read his logs, picked those three as good benchmarks, and proposed a comparison.
“Qwen 3.6 vs Gemma 4 E4B vs GPT-6 Luna”: three real tasks from Anand's own coding history, replayed from the commit just before his actual fix, each with a 15-minute limit. Open in a new tab →
The question behind it was practical: could a model small enough to run on his laptop in flight mode be useful for coding? The answer, on these three tasks and as of that day: GPT-6 Luna was clearly best on all of them. Qwen 3.6 was agentic — it explored, edited and tested — but ran out the 15-minute clock on two of the tasks. Gemma 4 “didn't really work.” “So that's going to be my default model when I'm on a flight and don't have access to models” was Qwen, with a shrug: “this obviously is something that will change over time.”
There was a subtler finding on the page that the talk didn't dwell on, and it belongs in this story. On the WhatsApp task, Gemma's run passed the existing tests while making no change at all. A green test suite is not the same as a finished task. So the benchmark treats the automatic verifier as an input and the actual diff — did it diagnose the right problem, make the smallest robust change, test it, and leave something he'd merge? — as the score. Hold that thought; it comes back.
“You don't necessarily need to hunt for a benchmark that suits your work. You can take your own logs and convert that to benchmarks.”
Anand, 52:28
Venky pushes back53:19How do you compare things you can't control?
Anand then invited the room to talk, and Venky took the floor with a challenge that was really two arguments. First: an agent is a model plus a harness, and you can't open up Claude Code and change what's inside it, whereas with open-source frameworks such as DeepAgents you write everything yourself. How do you compare things with such different control surfaces? Second, and more provocatively:
“So selecting models is becoming kind of obsolete. I mean, it's not a difficult decision to make: pick whatever is the cheapest and least complex, and then if you can write your own harness, put your efforts there.”
Venky, 53:19
Anand didn't defend the benchmark. He first checked he'd understood the challenge, then granted it: if model selection isn't a problem worth solving for you, “that's a perfectly reasonable decision.” And then he turned the question on himself, in a way that explained a lot about the ninety minutes.
“Anand, if you're happy with any of these models, then that means you don't have tough enough a problem that will be able to differentiate between models.”
Anand, quoting what he tells himself, 56:36 — his personal learning strategy, not advice for everyone
“I am trying to develop a sense of discomfort with that state wherever I can,” he said. “If agents are going to be doing all the work, what job will I have? I'm trying to discover that.” The discomfort is the method: look for tasks where one model succeeds and another fails, because the person who can write such a task is the person who can specify hard work for agents.
And how does he find those tasks? With a hilariously honest loop. Go to Claude or ChatGPT: “Boss, this is my problem, give me ideas.” It gives some rubbish ideas. “No, no, no, go research, look at my work, tell me something that I can relate to.” Then it says something he can't understand. “Simplify it.” And then: “A day, a week, a month later, something clicks.”
Find a problem that differentiates
One that fails under one configuration and passes under another. “The skill to do that, I think, will be relevant for at least a few more years.”
Write the verification criteria
The rubric — “even if it is a messy surface like an agent.” Long versus short system instructions; a fixed tool list versus auto-discovery. Take a guess, test the hypothesis, iterate.
Evaluate automatically
Pit A against B, one parameter at a time. If that isn't enough, let an agent keep tweaking the settings until the problem is solved reliably — “the equivalent of AutoML.”
“Specification, verification, execution is the chain that I'm trying to convert almost any product to.” He called the underlying skill turning “agent surfaces — that is, the kinds of things that I can change with agents — … into testable parameters,” and noted that it is what data scientists have always done with 100 messy parameters.
Act IV · The turn61:34Venky asks the question the room was thinking
Venky had one more request, and it was the one most of the developers were probably waiting for. “As agents become better and better at coding, what is the expectation from a software developer?” He'd like Anand to share what he'd found about coding in particular.
Anand said he would give the cynical answer first.
How to keep your job while the agents do it: three very real, very practical, very cynical suggestions
“Try and keep your manager in the dark as much as possible once they figure out that agents can do everything that you're doing and they're getting smarter.”
Anand, 62:10
- Security is your biggest strength. Whenever an agent might take over your job, ask: is it safe? Has compliance approved it? “What if it hallucinates? What risk percentage are you able to tolerate?” “Once you bring in the language of uncertainty, you will be able to kill many agentic initiatives. It is a very potent, very powerful weapon. Wield it well.”
- Generate output nobody will verify. Output with two properties: nobody is particularly going to check it, and if it's wrong it “either doesn't cause any harm or cannot even be detected.” Build a rubbish classifier — deploy not one but fifty — and if there isn't capacity to verify them, “there's a good chance that it'll go into production. And you've now deployed something to production, good for you.”
- Leave before the bill arrives. “Keep pushing stuff before the shit hits the roof, make sure you move to a different team or a different organization.” (“Don't wait for too long before hunting for your next job because you're building up a pile of risk that might explode at any point.”)
He then did the thing that turns a joke into an argument: he added that it wasn't only a joke. “Obviously, I'm giving this to you as anti-patterns… But it is useful also to think of these as strategies that people will naturally, even subconsciously, adopt.” If someone gives you ten tasks and you can blindly hand them to an agent and finish by midnight instead of 4:00 AM, “I would do that.”
Venky, quite reasonably, wanted the other half: “okay, so these are anti-patterns, but what are the positive approaches?” Anand: “Let's come to that.”
“It is very, very hard to verify output, and with agents, you can create a lot of output.”
Anand, 63:26. The cynical playbook only functions because generation has become cheap and fast while checking and owning the result has not. The rest of the session is about closing that gap — first by looking at what people actually do, then by building verification into the workflow.
Positive approach 166:00Your agent logs are a lab notebook
“Firstly, sharpening your tools.” If you use agents every day, you are generating a record of exactly what you do. Anand shared his screen and showed what he'd done with his own: he'd asked Codex to read through how he uses Codex. OpenAI keeps shipping features; which ones is he using, which is he ignoring, and what would it be worth?
The verdict, as Codex relayed it, was blunt enough that Anand's own reaction was to push back — “What rubbish! What does that even mean?” — before reading the evidence.
“Look Anand, when a new model comes, you have about 75 or 76% adoption. When there's a new workflow, your adoption is zero.”
Codex, reading Anand's own logs back to him — 66:30
New models: adopted quickly. New workflow features such as the status line and memory controls: “never used it.” The 11.3 hours saved that Codex's simulator quoted if about 40% of the gaps were closed is an estimate built on an assumption of about two minutes saved per resolved gap — not a measured result. The data story →
What he liked was that the analysis only counted a missed feature after it had shipped, so the gaps weren't false alarms. Parallel tool calls topped the list; second came cases where Codex could have asked him a question but his permissions blocked it. And his reaction to the savings estimate was entirely practical: “To save 11 hours I'm happy to spend two hours doing this. I'm even happy to spend five hours doing this.”
“The Hidden Gap In 903 Sessions” — the story Codex produced from Anand's logs, including the impact simulator. Open in a new tab →
He drew two lessons from it, both more general than the workflow details.
“Your agent logs are one of the most powerful diagnostic devices. Use agents on agent logs, and you will find a lot of stuff about your usage.”
Anand, 69:39 — the second lesson: you can't keep track of agent capabilities, and you don't need to, because agents can
The whole exercise compresses to a prompt you can steal for your own tool of choice:
“Find out all the new features of Gemini CLI [or pick your agent] that have been released in the last eight months, and find out which ones of those would have the maximum impact for me based on my logs.”
(Gemini CLI was no accident: as Venky said in his closing remarks, it was the tool the organizers hoped to have available a month later.) And then the question that took the idea from a single developer to a team: “What if you run this at a team level?”
Positive approach 270:36What 1,100 sessions say that no demo does
“Which incidentally is exactly what I did a few weeks ago for one of our teams.” Dharmendra leads a team of developers at Straive. Anand asked him to have the team put their Codex and Claude Code session logs in a shared Drive folder. Seven people did, and the folder held about 1,100 sessions — 1,112 top-level sessions after cleaning, spanning about a year, from machines on which none of Anand's own tooling was installed.
That last detail mattered more than Anand first realised. His opening instruction to ChatGPT was to prepare the data and find reusable skills and advice. Several steps into the work, he caught his own assumption:
“Hold on, these are logs from OTHER people's machines. They're not using my skills at all!”
Anand, to the agent doing the analysis — from the chat that produced the findings
Which reframed the recommendation: not adopt my skills, but a shared operating practice for the team, derived from what they actually did. And the agent's first full pass opened with a surprise of its own: “The most important finding is not what I expected: the team generally does inspect before editing; the weak point is validating what they finally changed.”
It is a striking picture. The agents and their users are doing the careful, exploratory part — reading code before they change it — most of the time. What they almost never do is go back and check the thing they finally changed. In Anand's phrasing in the room, more bluntly than the footnote would allow:
“Coding has become very fast, but testing and ownership has not followed.”
Anand, 71:26
He put the numbers plainly: “only about 2% of the Codex sessions and 3% of the Claude sessions had programmatic validation,” which in the room he summarised as people “giving it to you 97, 98% of the time without even a unit test.” (The careful reading is the one in the chart above.) “Standard basic practice would be at least test at the end, and even better practice would be write the test cases first and then do it.”
Why hadn't that been standard all along? Because for a human it is a lot of work. For an agent, “what difference does it make?” If it writes the tests first and codes against them, you are “much more confident that it's doing a good job. When you make an edit, regressions will happen a lot less often.” The same pattern shows up in Gemma's WhatsApp run earlier: tests that don't test the change are a false green.
The advice, which he called “by far the most common failure that I've seen amongst developers using agents,” was graded in three steps — each one cheap:
Show me the evidence
“At the very least, have the agent show evidence that the change worked, minimum, allowing you to manually test.”
Tell it how you would test
“Here is how I would have tested it; you test it the same way.” Backend: type-check. API: call it, then write test cases. UI: “open the browser, click, explore, run it.”
Failing test first
“Write tests first, or where possible, write failing tests first, and then implement” — a line for the global agents.md, “almost exactly the wording of my agents.md.”
In the Webex chat, Uttam had said the same thing a few minutes earlier: “ask the agents to do test-driven development.” Anand: “The biggest!”
Then, as he had every time, he put a boundary around it: “the general principle, again, is: don't rely on generic advice. This was based on experiments with a specific team based on the logs themselves.” The analysis produced a deliberately short list: after the last edit, run the smallest relevant check; if an edit fails because the target has changed, re-read before retrying; and for UI changes, finish with a real browser check. The recommendation was to try them on a few team members for a couple of weeks, then collect the logs again and see whether anything moved.
The point, in other words, was never the three rules. It was that the team had never been told what its own logs said — and that reading them took an agent and a handful of prompts, not a research programme.
The room asks75:03Open weights, regulators and the “is Jev GenAI?” problem
With the experiments done, the chat turned into a proper discussion. Two questions, in particular, showed how a room of bankers' engineers actually thinks about all this — and, once more, how Anand's answers refused to be universal.
How do frontier and open-weight models evolve in regulated industries?
Open-weight models lag the frontier by “maybe six to nine months”, as of the day, and Chinese providers have a strong incentive to keep releasing them. In regulated settings there can be “little choice” — in a US emergency room, no connected device is allowed at all, so the model must be deployed on site.
But “the path to production is pretty rare”: the quality gap and the effort are real, and the use cases where it's worth it are a small share. Three things must shrink: deployment friction (he now just tells Codex to install software and write instructions — “those are for a future avatar of you”), the quality gap (“Six months ago, I would not have coded on a flight. Now I am okay to code on a flight, within reasonable limits”), and perception bias, as with open-source software and security.
If Jev handles classification so cheaply, will classical ML in companies be revamped?
Anand expects more models in the middle ground between deterministic, structured ML and free-text generation — but doubts that this makes migration easier. People will ask, “Is Jev GenAI? … Will it make mistakes?” The answer “Even your ML classifier makes mistakes” meets “Ah, but no, that will hallucinate.” “Variety confuses more rather than less.”
The stronger force is the pitch that moves liability: “I will lower your cost by 80% and your risk by 70%.” The test for Jev, then: is it moving the risk-return frontier? “So far, it hasn't, at least on the few benchmarks that I've run.”
Then Venky asked whether there was more to present, or only questions. “Just one point, which won't take long.” The session ran past its end time — “I know we have spilled over, I will try and wrap up in just two minutes” — and the point turned out to be the one that completed the argument.
Act V · Reliability85:29If errors are independent, disagreement is a signal
The biggest problem with generative AI, he said, is hallucination: the model can be wrong, whether it is writing your code or producing output inside your code. Anand offered two clean ways to deal with it, and the first starts with an observation about how models fail. In a benchmark of 11 models classifying customer-support messages (when will I receive my order? — delivery-period queue or track-order?), some questions were easy for everybody, some were hard for everybody, and some tripped only some models. The mistakes were not strongly correlated — one pair he showed correlated at about 70%, another at only 5–10%. “Meaning they are making independent mistakes, which means I can double-check, cross-check.”
Guess before you scrollOne model averaged a 14% error rate. If you only let an answer through when two independent models agree, and send disagreements to a human, what happens to the error rate — and how much work does the human get? What about five models?Click to reveal →
Two models: error falls from 14.1% to 3.7%, with humans reviewing 12.6% of cases. Three models: 2.2% error, 18.5% review. Five: 0.7% error, 28.1% review.
“You may not be able to get rid of error, but the cost of models is not that high.” Those figures are for this task and these models — the benchmark page itself warns to drop models that are consistently bad.
“Dealing with Hallucinations”: 11 models, a customer-support intent task, and what happens to the error rate when more models must agree. Open in a new tab →
Then he gave the operations manager's version. “Look, I will reduce 72% of your work and give it to you at 99.3% quality.” “Hey boss, my team only gets to 95% quality. 99.3 is way better than what I can do myself with humans, and at almost one-fourth the cost.” The framing matters. Verification here isn't a gate at the end. It is a design that routes uncertainty to the right place.
The second way88:51Ask the model how sure it is — then don't believe it
The other approach: ask models for their confidence. Logprobs are one way; the model's own stated percentage is another. But, he warned, when a model says it is 90% confident, it is not right 90% of the time. In the example he showed, GPT-4.1 Nano's “80 to 90% confident” answers were right only about half the time. “Overconfidence is consistent across models.”
“You can calibrate that. And you can say, when something reports 85%, treat that as like 50%.”
Anand, 89:30 — build the calibration curve from known answers, then cross-check only what falls below your cutoff
The linked calibration study carries the same method to a larger set of 770 banking requests and spells out the recipe: ask for uncertainty in a better way, replay reviewed cases with known answers, compare risk signals (logprobs among them), choose a cutoff from your acceptable error rate, freeze it, and confirm it on untouched cases before auto-passing anything. That last step matters. According to the project's report, a threshold frozen on the 770 development cases auto-passed 65.2% of 2,310 untouched requests at 4.38% observed error against a 5% target — but a looser 10% rule frozen the same way missed its target. Pass or fail, it was tested on cases it hadn't seen.
“Does 95% confidence really mean 95% correct?” — 770 BANKING77 requests, several confidence signals, and a frozen threshold tested on held-out cases. Results are specific to this task and these models. Open in a new tab →
And for code, the same thought, applied to the problem the team logs had exposed: “Have an agent adversarially test it. Find all the mistakes in this.” That is what a code review is, he pointed out, and there is no reason a human should be the only reviewer. “Let agents talk to other agents, and find all the mistakes, and then what comes to you should be something easy to review: ‘These are the things I'm not sure about.’” The person doing that last review “needs to be really experienced, but that's what we're working towards.”
Epilogue90:31The ending was already written
The last minute of the recording is unusually tidy. Anand had spent ninety minutes treating other people's advice as hypotheses; now he applied the same treatment to his own. “Look, people will say lots of things, I will happily say all kinds of things. These are things that might have worked for me.” Then he ran through the one test he'd stand behind.
“You should just test. Whether it's prompting advice, what model to use, what agents to use: test it.”
“Your logs—session logs, and any logs in general—are probably the best way to improve any kind of workflow.”
“Make sure verification is part of that workflow. Use test-driven development, use double-checking or triple-checking as a means for validation, use confidence calibration.”
“Benchmarking is not just for your workflow; benchmarking is a verification mechanism and should be literally part of the delivered workflow as well.”
— Anand, 91:30
It is worth noticing what that ending doesn't say. It doesn't say use this model or write prompts this way — every specific recommendation of the session had a date and a dataset attached. It says that the method is the durable thing, and that the method has two halves. Use experiments to find out what's true for you. And build the same kind of experiment into whatever you ship, so the thing that checks your work is part of the work.
Venky closed with the part of the story that no benchmark would have predicted. “Unfortunately, a lot of tools that you have showed, we cannot access from office; maybe we'll need to do it in our personal capacity.” The hands-on workshop Anand had originally planned couldn't happen because the software wasn't available — but “probably a month from now we'll have what you wanted in terms of Gemini CLI,” and perhaps one more session. For a team limited to a single agent, the empirical habit is arguably worth more, not less: you may not be able to choose your tools, but you can still test how you use them.
Anand's 70:13 prompt asks which new features of your agent would have helped, based on your logs. The team-log result suggests a second question worth the same treatment. Export a week or a month of your agent logs, give them to an agent, and ask it:
What am I repeatedly doing badly? Which new features of this agent would have helped? What one measurable behavior should I change?
Then do the part the talk kept insisting on: check the recommendation against your next set of sessions. If the number didn't move, the advice was a hypothesis, and now you know.
Six things to take away
Every result is specific to its task, its models and the date of the talk — 1 October 2026.
