Sixteen people answered a five-question form at the start. Two hours later they were browsing an app built out of their own answers — and the app had told Anand something he didn't know he wanted.
The full two hours, compressed with AV1 from 900 MB down to 130 MB — 2 h 09 m of video at roughly a megabyte a minute.
📄 Read the transcript ·
🎧 Download audio ·
🎨 The whole session as one comic page
Earlier in the series: 1 · Context Engineering ·
2 · Tools & Workflows ·
3 · Agentic Analysis
The last of four workshops opened with a sentence that seemed designed to make the topic disappear.
"The bulk of today's session will be about coding, but just as much about how we don't even need to code."
Anand, first words of the session
Then the room display refused to cooperate, the screen share went recursive, and Anand found himself looking at his own desktop reflected inside itself. "My desktop was actually a thousand desktops out here, one inside each other." Saurabh, hosting, called it correctly: "So like Inception, right?"
Once the mirrors were untangled, the session went where it had been going all along — not to a code editor, but to a scatter plot from 2008.
Anand's answer to "why write code at all?" is not software. It's a picture of every movie ever made — 15,510 of them, in fact. Popularity on the X axis, rating on the Y, one box per title. Inception far right (2.86 million votes). Band of Brothers near the ceiling at 9.4. Twilight "hovering around five" — 5.4, to be exact. The Hunger Games somewhere in the crowded middle.
He pointed at a box in the bottom-right corner — the impossible quadrant. Ridiculously popular, rated terribly. "Any guesses on what this might be?"
The room guessed Shawshank Redemption (wrong direction entirely — 9.3), then Despicable Me, then Animal. Nobody got it.
The answer was Snow White (2025) — Disney's live-action remake. 2.2 stars. 397,875 votes. In the entire dataset, no film rated below 2.5 has even half as many votes — Radhe (1.8, 181,897) is the runner-up, with Sadak 2 (1.2, 97,354) further down the same lonely wall.
"I love these outliers."
Anand
But the outliers are only half of it. The same chart, filtered to Animation and dragged through the decades, becomes a history lesson.
In the 1930s there is exactly one animated title in the whole dataset: Snow White and the Seven Dwarfs (1937), rated 7.6. In the 1940s a cluster appears, led by Pinocchio, Fantasia, Bambi and Dumbo. The 1960s is when Studio Ghibli starts climbing toward Disney. And by the 1980s — the "biggish explosion", as Anand called it — the three highest-rated animated titles of the decade are Dragon Ball Z (8.8), The Simpsons (8.6) and Grave of the Fireflies (8.5). Not one of them is Disney. My Neighbor Totoro (8.1) sits above The Little Mermaid (7.6). Then The Lion King and Toy Story flip the story again.
Which leaves the chart with a joke buried in it that nobody in the room noticed. The film that opens the entire history of animation, and the film sitting alone in the disaster corner, are the same title, 88 years apart. 7.6, then 2.2.
The other end of the chart is the one Anand actually uses, though — the quiet, highly-rated titles nobody has heard of. He pointed at Steel Ball Run (9.5), Sapne Vs Everyone (9.2), Cosmos and Planet Earth II (9.4): "A good way to discover things that are very highly rated but we probably haven't heard of."
Play with it yourself. Every dot is a movie; the Outliers checkbox is the button that took eighteen years to exist.
sanand0.github.io/imdb — filtered to Animation, as Anand showed it. Tick Outliers and drag the year slider.
Built in 2008 after Col Needham, IMDb's founder, noticed someone scraping his website and invited him to Bristol for a chat.
That origin story is the good bit. Anand had been scraping IMDb; Needham's response was "why don't you come out to Bristol and let's have a chat" — and the visualization they ideated together, he says, kicked off his entire data-visualization career. Eighteen years later it's still running.
And still improving, which is the actual point:
"I've been looking at this, I never thought of adding a button for outliers. Not thought of, it was a reasonable amount of work. When AI-made coding becomes easier, it was just a, 'Oh, I've been giving this for so many years, just do it for me.' 'No, no, no, that's not how I want the outliers, do it differently.' 'Okay, yeah, this looks fine.'"
Anand, on the feature that waited fifteen years
The value got named out loud in the room, in one line: "typically what happens is when you have a top 10, top 20, you tend to go and look at those 20; your decision gets restricted to that. That outliers setup is a very beautiful way to look at things which you probably like."
Asked how hard this would be to reproduce, Anand didn't hedge: "Download the IMDb data and create a scatter plot matrix of rating versus votes. You'll be between 50 to 70% there."
Debi named the category before Anand did — "So like a dashboard for yourself." Anand corrected the plural and kept it: "Dashboards for ourselves."
"Personally, I find this sort of a decisioning interface as one of the most powerful reasons for using code… Often we need the code, or we need to explicitly get it to do things that may involve code, because we want to play around with the code in a certain way, and a chat interface doesn't do the job for us."
Anand
Then came the move that made the rest of the session possible. Anand put up a QR code and asked everyone to fill in a short form — not as a warm-up, but as raw material. "What we're going to do is run the rest of this session to build a decision interface from your responses. And effectively we are feeding the data for the system to get built."
The questions were disarmingly personal for a corporate AI workshop:
Sixteen people answered. Here is what the room actually looked like, in numbers.
The homework from the previous session split the room in a telling way: 9 of 16 had automated an annoying task, 7 had gotten an agent to do something for them with Python, but only 4 had managed to make it fail. Two people listed Hugging Face and Perplexity as their coding agent, which Anand noted with visible delight.
Every chart here is drawn from the same CSV Anand handed the agent at the start of the session — the one that became the Room Map.
Topics people offered help with vs. topics people wanted help with · 16 respondents
The room could not help itself with the thing it came for. Career & Future of Work drew four requests and zero offers; AI & Coding drew four requests and one offer — Anand's. What the room had in surplus was cocktails, hikes, Netflix picks and Singapore trivia. The one topic in genuine balance was money: four offers, two requests, and five of the eight "decisions I keep repeating" answers.
| Topic | Offered | Needed |
|---|---|---|
| Lifestyle & Travel | 5 | 2 |
| Investing & Markets | 4 | 2 |
| AI & Coding | 1 | 4 |
| Career & Future of Work | 0 | 4 |
| Business & Research | 1 | 3 |
| Family, Health & Education | 1 | 2 |
| Productivity & Organization | 1 | 1 |
| Relationships & Networking | 2 | 0 |
| Trust, Safety & Meta | 1 | 0 |
Free text, so people naming two are counted twice
Claude Code by a distance — and two people who answered Hugging Face and Perplexity, which Anand noted with visible delight. A quarter of the room was not using one at all.
Three exercises, 16 respondents
Breaking it was harder than using it. The gap between 9 and 4 is what launched the session's first real argument — "what constitutes failure?"
The failure gap became the session's first real idea. Making it fail turned out to be the hard homework. "'Make it fail' seems to have been either tougher or less attempted," Anand observed. "Is that because we don't try too hard, or it's too capable? Both are possible."
Debi pushed back usefully: people had probably interpreted the exercise differently. "In my case, I think I could make it fail because I wasn't getting my required solution. So it was just failing, and then I said, 'Okay, what is happening?'"
"I think it's an important point, which is what constitutes failure? … just knowing that 'here's something that it cannot do as of date' which on a future day if we're able to do, literally represents the boundaries of it."
Anand
He keeps a running list of these — an "AI bottlenecks list", updated daily, of things he tried and failed to do. And the entries have started to change character in a way that is either funny or ominous depending on your mood:
"Increasingly, the list comprises of things like: I have too much work to do because of AI. I'm supposed to sit and read this response and action it; that is a lot of work for me. But it is still a bottleneck."
Anand
"You don't necessarily need to think of it as code versus not code. That used to be a distinction pre-AI; it no longer is a relevant distinction."
AnandTo prove the point, Anand ran the most trivial program imaginable — twice.
Write and run a Python program to print the 50th Fibonacci number.
Deliberately on the lightest, cheapest models available.
Both wrote it. Both ran it. Saurabh, reasonably, asked whether you needed Codex for this. "This is normal ChatGPT."
The point of a trivial example is that it doesn't take much by way of intelligence. ChatGPT and Claude each have "a little computer sitting inside" — a container that can write code, run code, and download packages. So why does Claude Code exist at all? Why Codex?
"Codex and Claude Code are effectively ways of getting the intelligence of these models into your permission systems. When you run something on ChatGPT or Claude, the permissions it has are the permissions that it has. When you run Codex or Claude Code, the permissions you have are the permissions you can give it."
Anand — the sentence the whole session hangs on
Not the data on your machine, though that matters. Not the compute. Mostly the access. Your laptop is already logged into your company's email, OneDrive, SharePoint, Salesforce, HubSpot, the ERP. That access is the asset.
Anand cut through the recent naming churn with unusual bluntness. Both vendors have shipped a mode called Work, and it confused the room until he said this:
"For all practical purposes, there is no difference between Claude Work and Claude Code. They are the same thing… When somebody says 'Work,' think of it as 'Code,' but with a lighter name so that people are not put off by it. It is really more marketing than anything else, making things easier for us also — nothing wrong."
Anand
The distinction that does matter is a one-liner, and Debi got there first: "it's basically… using it on the browser versus your downloaded app." Browser has no connection to your local data. Local app roughly inherits your permissions. (And remote control has since muddied even that — you can now run it locally and drive it from the browser. Anand was refreshingly honest: "I haven't really gotten a good mental model around it.")
So he tested it live, with the dumbest possible question — "How large is my disk?" — because a cloud container physically cannot answer it.
| Where it ran | What it could see | Result |
|---|---|---|
| ChatGPT desktop → Chat | Cloud container | "I can't inspect your computer's disk from this chat." |
| ChatGPT desktop → Work / Codex, full access | The actual machine | 937 GB total |
| Claude desktop → Chat and Work | Cloud container | Reported the session sandbox, not the laptop |
| Claude Code, given a folder | The actual machine | Correct answer, after one permission prompt |
This is the part a slide deck can't do. Sandeep, on Teams, ran the same prompt and reported back: "Claude basically told me, 'I have no way of checking your actual laptop disk.' It gave me the size of the session sandbox." Anand assumed Chat; Sandeep clarified he'd tried Work too. Anand signed in and reproduced it. Another participant reported 252 GB with 30 GB free — from the cloud sandbox, not their laptop.
Someone else, poking at the same answer, discovered their SSD was made by SK Hynix. Which prompted Anand's aside about what code is actually most used for:
"This, incidentally, is one of the most common uses of code, which is: 'I have a problem with my computer, help me fix it.' I had a problem with the NVIDIA driver — fix it. My computer is running slowly — fix it."
Anand
A participant immediately objected that "My computer is running slowly, fix it" might go and delete a lot of things you didn't want deleted. Anand's fix was three words plus a phrase: "Tell it 'fix it safely,'" or "Please display all the steps you are going to follow," or "Ask me first before you take action."
Which produced the best laugh of the hour, from a participant watching himself negotiate with a machine:
"Increasingly, I'm feeling like someone of my father's generation who would ask me, 'Look, do this,' and before I do something, 'Are you sure? Don't touch the bonnet!' I know, I know, I kind of know this stuff. I do this for a living. And these are far more patient than I am. Clearly."
A participant
And the best story, which was really an argument for local access all by itself:
"I take a lot of photos and I have this iPhone since 2009 or something. So I tried to connect it with the camera and I couldn't. Apparently, there is a Bluetooth stack of devices from 2009 which has overflowed. So I had to remove it, and I don't think I could have figured that out through Google or anything like that. I was just blown away by this."
A participant
Anand R. — a different Anand, in the room, and the association's secretary — asked the question everyone was carrying:
"You're trying to check out how much memory you have, so you gave it access to the computer. Now, because of the nature of the task, it's going to check every folder and file. What is the risk that you're now running of some unintended thing coming in?"
Anand R., IIM Alumni Singapore
The answer came in three parts.
His evidence was a Simon Willison interview with the Claude Code team, where they said that almost everyone inside Anthropic runs auto mode. Anthropic then made it the default, and documented it as a permission mode.
"What they're saying is a team as sensitive as Anthropic is comfortable just running code without any sandboxes and any kind of checks using auto mode, because every instruction, every tool call that is being made is being checked by a model as smart as Sonnet 3.5 to see if it is okay in this context or not."
Anand
"Is there a risk today, therefore? The answer is negligible."
Sonal asked the follow-up that actually decides adoption: "if I just asked it to not just find the code but also just run it, what is the risk in terms of let's say if it does it wrongly? Are we able to undo what the code does?"
Anand split it in two, and the practical half is the better half:
"The practical answer is: it doesn't matter what the technology says or what I say; what matters is, has it gained enough of your trust that you would allow it to do something? Until it does, make backups. Copy that folder somewhere else… Try it five times, ten times. Copying takes very little effort. And after every single one of those attempts, if you see not even a sniff of anything going wrong, then the next time you may tell it, 'You make a copy, make sure I can undo.' And then after a month, a year, however long it takes for you to get comfortable, don't bother telling it to make a copy; it will fade away."
Anand
Sonal's second question — can agents be put on a scheduler? — got a one-word answer: absolutely. Both vendors have schedules; in Claude they're called routines, in ChatGPT scheduled tasks. The only catch is that your machine has to be awake.
And when a participant asked whether open-weights models would be safer, since people can inspect them, Anand declined the easy answer:
"They're all black boxes. Not only are they black boxes to us, they are black boxes to most of the experts as well… Anthropic kind of has something like 'brain probes'… The comfort, if at all, is the other way, which is: can I hold somebody liable?"
Anand
A participant selling enterprise AI services confirmed the mess from the buyer's side — security questionnaires, unclear liability, nobody with an answer. Sandeep's version was funnier: "It's those 78 questions once, then I answered those then I got another 65." Anand's verdict on the whole category: "I feel the questions I'm answering, they're not even the right questions they should be asking." Too early. "The cleanest thing is to protect the blast radius."
The second use case is the least glamorous and probably the most immediately useful: pointing an agent at the mess on your own disk.
Anand dictated the task live, in one breath, to ChatGPT's Work mode on a deliberately light model:
"I'd like you to go through my photos folder and tell me how large it is. It contains a bunch of photos and videos. I'm interested in reducing the file size without reducing the visual quality and also preserving the metadata. Research what are the best image formats that I should consider for this and test on a few folders to see how much compression I can get. Extrapolate this to the rest of the folders and give me a sense of, maybe on a folder-to-folder basis, how much I can expect in terms of compression and show this to me in a visual form."
Note what's in there: research the formats, test on a sample, extrapolate, then visualise. That's a work plan, not a command.
First pass came back with ~4–5 GB of savings on images and a surprise — the videos, not the photos, were eating the disk. So he pushed:
"Video compression, actually, is now shockingly advanced. AV1 is a phenomenal, phenomenally powerful format."
Anand
If you want the proof, it is already on this page. The recording at the top is an AV1 file: 2 hours 9 minutes of screen share and camera video, 900 MB compressed down to 130 MB, with nothing visibly lost. Seven times smaller, for the cost of letting a machine grind overnight.
Second pass: 15.6 GB of savings, plus a side-by-side comparison of AVIF quality levels — the agent doing what you would do, which is shell out to ffmpeg and read the file sizes back. Anand's reaction to the comparison grid is the most quotable sentence about image compression ever spoken in a workshop:
"I can't tell the difference between any of these, frankly. Yeah, and that's usually the case. These are way too conservative and I'm quite happy to be more aggressive."
Anand
The economics matter here, and they're counterintuitive. Running a 70 GB compression job overnight costs almost nothing in tokens, because the tokens are spent writing the instructions, not doing the work. "When the program is running overnight, the tokens are not getting consumed other than maybe to monitor it a little bit… These are not complicated tasks. They have to write simple programs, they have to delegate the compression to a powerful compressor."
The travel spreadsheet was assembled from emails, flight tickets and scanned passport images. Debi identified the use case instantly: "How many days are you in Singapore versus India?" Anand's version has a legal twist:
"There's apparently a Karnataka High Court judgment that says that the day you enter India does not count, the day you leave India counts. So depending on this, I'm either in India for 117 days or 123 days, and that obviously changes the tax treatment and all of that."
Anand — the judgment is on Indian Kanoon
The music library is the better story. Thirty years of MP3s — 1,412 files, 103 hours — with essentially no metadata beyond a filename convention Anand had maintained by hand: movie.songname.mp3. No language, no year, no official title, no composer.
So the agent went and got them: Wikipedia for the songs, MusicBrainz for the catalogue (album IDs and recording IDs are separate, and it matched both), and a plain web search for the leftovers. The results are worth looking at, because they show exactly where an agent's reach ends.
A participant asked the sharpest question of that stretch: "Where does this code stay?"
"This code is ephemeral. The code is not saved anywhere. It writes the code, it runs the code, and gone. It doesn't even save it."
Anand
Which prompts the obvious objection — why are we wasting tokens to write the code again and again? Anand's rule of thumb is refreshingly unromantic: he runs the photo job once every two months, so he doesn't care. If he ran it every two days, he would.
The music job crossed that line, so the code got saved. It is now musictag.py — a 413-line Python script with dump, fix, check and clean subcommands, a CSV as the system of record, and MusicBrainz IDs written into the ID3 tags via mutagen. Anand runs it manually because he's comfortable doing that. You don't have to be.
"Save the script because I'm going to ask you again tomorrow to do the same thing." Next day: "I've added a whole bunch of songs here, you run the script, make sure you know where you've saved it, or I'll tell you where you've saved it, and run."
"But the point is, the code is not for us to run; the code is for the agents to run."
Anand
Anand started with 1,412 MP3 files and one piece of information: a filename he had typed himself, movie.song-name.mp3. Everything below was found by the agent, and is now the CSV that musictag.py keeps in sync.
Share of 1,412 files carrying each tag · striped bar = the only field that existed before
It matched two thirds of the collection to a MusicBrainz release ID and 61% to a Wikipedia page — for songs whose only identifier was a filename Anand had typed by hand over three decades. Lyricist credits are the field that stayed sparse: 178 files. That is what an unsolved problem looks like in a chart.
Songs per decade · 103 hours of audio in total
Of 1,286 songs with a composer credit
Ilaiyaraaja and A. R. Rahman account for a third of the credited collection between them — a fact that did not exist anywhere until an agent went and looked it up, song by song.
Debi asked about the Gmail connector. A participant answered before Anand could: "It's fantastic, Debi. It's fantastic, please use it."
Anand's advice on connectors versus plugins versus whatever they're called this quarter:
"If you're able to access your email in any way, don't worry about plugin versus non-plugin. It simply means that at some time you would have enabled it. It used to be called connectors, now it's renamed to plugins or whatever."
Anand
And then Sonal supplied the session's single best failure report — a genuine, current, reproducible boundary:
"I have like hundreds of thousands of unread emails — it could not actually delete emails. So it was able to sort of work with me to move emails to trash, but the actual physical deletion of emails had to be done by me. So that's still, I think, what it's not capable of doing… I'm down to like 60,000 unread emails from 130, 140."
Sonal Priyanka
Anand's response: "That's probably good, actually." Once it's in the trash, your job is done.
From the live form: "a decision you repeatedly make using too many tabs, sheets, notes, or people." Eight answers came in — and five of them were some flavour of which stock, fund or company to bet on, while three people in the same room list investing as something they can help with.
Someone at IIT Madras had asked Anand to update his CV so they could consider him for a professor-of-practice role. His CV was a PowerPoint deck from 2024, still describing him as CEO of Gramener, a company that had since been acquired by Straive.
He knew exactly how not to do it:
"Because I know it can edit HTML well, but not Word or PowerPoint. So for the last year, my strategy has been: don't try and tell it to edit PowerPoint decks. Convert the PowerPoint to what it can edit, and stay in that space."
Anand, answering Sonal's "why HTML?"
There was a second sneaky move underneath it. The whole conversation ran through what Anand calls local MCP — "Think of local MCP as where I'm giving ChatGPT a connector or a plugin to my computer." (Model Context Protocol is the open standard behind it; it's what has since turned into "Remote Control" in these tools.) His stated reason is gloriously mercenary: "for all practical purposes, ChatGPT's tokens are free — Claude's are not."
So step one was not "update my CV." Step one was "help me convert this into HTML" — and then, crucially, a quality bar rather than a spec:
"Take this and convert it to HTML, and I won't be happy until you've gotten to something really close to the original."
Here they are, side by side — the 2024 PowerPoint-turned-PDF, and the 2026 HTML that descended from it. Anand's own summary of the conversion: "I looked at this one, I looked at this one, and I couldn't really tell much of a difference."
Both are live HTML pages, scaled to fit — the same fixed 540×780pt sheet the PowerPoint used, which is why they print back to a near-identical PDF.
Watch the gold medals, the school prizes and the Gramener client logos disappear. That is "this is a kid's CV", executed.
The whole redesign conversation: "Anand CV Update 2026" on ChatGPT · read the export here
Two things about that conversion are worth stealing.
First, the honest accounting of how much steering it needed:
"In short, if it had been allowed to do it by itself, it would have gotten there maybe 80%. If I had blindly — without looking at the output — said 'do better' three times, it would have matched it close to pixel-perfect."
Anand
Second, the hidden machinery. To compare its HTML against the original, the agent wrote code that drove a headless browser with Playwright and printed the page to PDF — without Anand ever opening a browser. "Browser automation is a very powerful capability that the systems have."
"When editing, the broad rule of thumb is: we are able to convert between formats far more seamlessly than we thought before. Shift, edit, shift back."
Anand
Sonal wanted to see the actual result and asked the interesting question: did it just cut words to fit? Anand's answer is the part of the story that has nothing to do with technology.
"It basically ChatGPT said, 'Look, this is a kid's CV. As an adult, you should be talking about adult stuff. The fact that you won a gold medal in school is charming, but please take that out. Talk about what you are currently doing and what you can do.'"
Anand, relaying the editorial verdict
The chat export bears this out — and shows the model refusing the brief it was given. Anand asked for an update. It proposed a repositioning.
"The biggest change I would make is not 'update the 2024 CV.' I would change what the page is about."
"The existing page says, roughly, 'I have had an unusually strong 30-year data/consulting career.' The 2026 page should say 'I am now doing unusual, hands-on work at the frontier of AI — in industry, education, research, and public communication — backed by that 30-year career.'"
"I would preserve the one-page 540×780 pt layout, Constantia/Calibri typography, approximate font sizes, two-column structure, four-image strip, and timeline. But I would replace perhaps 60–70% of the words."
It also did something more interesting than writing copy. Asked to pick portfolio images from Anand's own repositories, it rendered the candidates as contact sheets and judged them at the size they'd actually appear — roughly a 68-point-tall thumbnail:
"I would not use the multi-LLM double-checking chart or system-override matrix: excellent work, but they lose visual impact when shrunk."
"I would not put a new headshot at the bottom. The current visual language is 'here is evidence of what I make,' which is unusual and much stronger. A headshot would turn the page toward an executive bio."
This is the Anthropic-versus-OpenAI temperament question in miniature, and Anand came back to it later. But note what made it possible: the model had to be able to read his disk. "Go through my entire disk. This is what I was two years ago, you figure out what I am now and put that in."
A participant asked how this differed from just using chat, since chat had happily edited their Word file. Anand's answer was unusually candid about his own uncertainty — and then landed the point anyway:
"However, I am only about 30% sure of this. In fact, I am not even sure if I did this on Codex or ChatGPT. Like I said, I have a connector from ChatGPT to my local computer… And therefore, does it make a difference? Not really. So, we're using code, we're giving it access to local data, and that access is what makes a difference."
Anand
A participant described how their team escapes corporate AI slop in slide decks — and it's a good recipe:
"In the corporate world, there's so much of AI slop, you can figure out which one is which — but the most sophisticated version of doing it is: if your corporate has a color schema, you kind of tell Claude, 'This is the color schema.' … After a few iterations, I've got to a point where my slides look like I've made them. But it has imperfections. I don't want perfect slides."
A participant
Anand's addition: "That style transfer makes a big difference when we provide examples." And when copying a style, asking for a checklist to verify against turns out to be more powerful than describing the style again.
Anand R. offered the other half — what happens when you give it your whole corpus instead of one document:
"I thought, why not make a website about myself? … I downloaded LinkedIn data and fed it in… It has a phenomenal amount of detail on what I've done in life, or what I have views on, that even I did not remember, which might be very relevant today if I wanted to position myself for something. … I've lived in many cities and managed people in many cities, so I kind of got them all in a graph. So that became my favorite tool when I was talking with people."
Anand R., IIM Alumni Singapore
His advice to Sonal, unprompted and rather good: keep one folder where every CV, bio and note about yourself accumulates. Over time the agent stops answering questions and starts offering advice you didn't ask for.
Now the promise from the start of the session came due. Sixteen people's answers were sitting in a CSV. Anand opened a fresh session, picked a medium-intelligence model with full access, and started dictating.
What he dictated is remarkable mostly for how much it admits:
"I'd like to build a networking application of sorts based on the responses that we got on the survey. The aim is to connect people with each other. To be fair, I'm not really sure what kind of utility we could provide… It's really more a decision intelligence app…
I know this is vague, but you are supposed to be smart. Use your intelligence. Create something that is both useful, practical, and simple. Put it as a web interface. I would like to be able to literally email this to people, so maybe just a plain HTML application with the data embedded in it would work fine.
And since people won't even know what this is if they open it blindly — that includes me — I'd like the application to be reasonably self-explanatory. Don't go rambling about what to do etc.; the whole point is for you to create this in a way that is both intuitive and useful."
Anand deliberately switched model families, and explained why with reference to EQ-Bench — a benchmark that scores models on traits like compliant, challenging, warmth and validating alongside raw ability.
"The model set of models that challenge you the most are the Anthropic models. They say, 'Look, you are saying X, I actually think Y is probably what you want, I'm going to nudge you towards Y or do Y.' When I'm less clear, I go to the Anthropic models. When I'm more clear, I go to the GPT models. When I want to feel good, I go to the Gemini models."
Anand — the most quotable model-selection heuristic of the session
The Gemini characterisation came with a demonstration: tell one "no you are wrong, 2+2 is 1" and it will find a way to agree with you.
While the agent was working, Anand typed another instruction and hit enter. It didn't stop. It read the new line mid-flight and carried on.
"This is called steering. That is pretty powerful — because if it's in the middle of doing something that is taking two, three minutes, then you realize, 'Oh wait, I forgot to say this.' You don't need to stop it; you can just tell it whatever else, press enter, it will take it up."
Anand
Later, doing the same thing in Codex, he noted the behavioural difference: Claude Code absorbed the instruction into the current run; Codex queued it for the next turn. Both work. Neither is in most people's mental model of "chatting with an AI."
Anand walked through the four ways to ship software — desktop executable, mobile app, web app with a backend, web app without one — and then made a claim that explains most of what he builds:
"So the magic word to use in many cases is HTML application. Single-page HTML application makes it even more precise. I want one HTML file which is the entire application. Then you can mail it, you can share a link."
Anand
His supporting argument is worth internalising: a huge fraction of what we do in Excel needs no server at all. You only need a backend when you want to save something. Gmail reads offline. News apps read offline. An interface over data you already have is a file, not a service.
Then, to publish it, he did something that should not work and does:
"I want to publish this on the public internet. There are supposed to be lots of servers that allow agents to publish HTML files. Find and push."
A participant asked the obvious sceptical question — why would anyone provide that infrastructure for free? Anand's answer was the biggest laugh of the session, and also completely correct:
"It costs nothing. You are putting your data into my server, which I can happily do whatever I want with it, at the very least read. If you have 2,000 of these, I will sell you as my asset and get acquired."
Anand
The app opened with a headline nobody had asked for: 62% of the decisions people repeat are about money — and three people in this room already do that for a living. It had found supply and demand sitting in the same room and pointed at the gap.
Here it is. It's a single HTML file with the data embedded, exactly as specified. Click any person to see who they should talk to.
The Room Map — built live, from 16 form responses, in a single self-contained HTML file. Names appear only for the 8 people who consented; the rest are matched anonymously.
Anand read the matches out loud as he found them. "Anand for real AI use cases, not quick answers; Anonymous 5 for…" — then, catching himself, "Okay, that's whom I could help. Rohit: ask me how to use code, tell me every bit of everything."
And: "The examples are so detailed. And I can help — obviously Rohit, but you could also help Shijo with investment advice and Anonymous 1 on cocktail and AI investing."
Then came the sentence that justifies the entire exercise. He looked at the finished app and said it wasn't what he wanted.
"Now this is not what I need, but I didn't know what I needed before I asked it in the first place. Now I know — and therefore prototyping becomes a very powerful way of need discovery."
Anand
Debi named it: "It's created a decisioning interface, sort of a thing." Exactly the category the session opened with, arrived at from the opposite direction.
So he dictated the revision — "Modify the application so that you show me a list of all the names, anonymous or actual, and when I click on each person, it should show me who are the people that person should connect with" — and pushed again. That version is the one embedded above.
One form field read: "If you could ask the others here one question and get an honest answer, what would you ask?" Shown without names, as promised. Read them as a group and a mood emerges.
Code has one property that chat does not: it is deterministic. Anand's fourth use case exploits exactly that — take a document whose whole purpose is to encode rules, and convert it into rules that actually execute.
His example was an insurance logic engine that compiles a motor insurance contract into something halfway between English and Prolog:
A valid claim is true if the driver is fully eligible, and the vehicle authorization is valid, and the incident circumstances are covered, and the claim procedure is followed.
The driver is fully eligible if they are age-eligible, they are license-compliant, and the driving history is clean.
They are age-eligible if age > 21 and driving_experience_months > 12.
"Goes on almost like a tree," said Anand — and that tree is the claim process. Feed in a claim: Marcia, 28, 84 months of driving experience. Age valid. Licence valid. Then one branch fails. Blood alcohol must be under 0.08; hers is 0.14. Denied.
InsurLE — Insurance Logic Engine: 4 contracts, 8 claims, 48 logic rules, and an animated trace of every branch of every decision.
Related: Policy as Code — upload any policy PDF, extract atomic rules, then validate documents against them, entirely in your browser.
The obvious objection — what if it converted the contract wrong? — Anand answered by counting how many times you have to worry:
"One may argue that it did not convert it correctly to the program; the contract needs to be verified and matched against the program. That's a one-time effort, and you can do this very diligently with accuracy. The second is whether the claim is not converted. Okay, cross-check, double-check, triple-check four times, whatever. But once those two are sorted, there is no arguing whether it's done the job right or not."
Anand
"Upload any arbitrary document and say, 'What part of this can be converted into programmatic verification?'"
Anand — the reusable move
A participant asked for the distinction between programmatic and LLM verification. The answer is short:
"Programmatic verification is deterministic. An LLM verification may not be. … 99.999% of the time, it will likely be correct — yes or no, yes. And 99% of the time it will give you the same answer. It may simply miss the blood alcohol count once."
Anand
What followed was the best discussion of the session — the room arguing, correctly, that determinism isn't always what you want. Anand R. pointed out that an elevated blood alcohol reading "could be because of some medicines or something like that." Sandeep took it further:
"There may be situations where… on compassionate grounds, what do we do kind of a thing, right? So rather than saying a digital yes or a no, what are the other circumstances that should be taken into account and assist the decision-making?"
Sandeep
Anand's resolution didn't pick a side. "That can be brought as another overlay also onto the deterministic." Rules decide what's decidable; judgment sits on top of it, visible and separable — which is precisely what the comic page draws in panel 6.
"I think determinism is good wherever we can afford it."
Anand
The same trick works one rung down. Take the European financial promotion guidelines — a stack of PDFs governing how anyone may advertise a financial product in Europe (in the UK, the FCA's COBS 4 is the canonical version of this genre). Step one isn't code. Step one is: use an LLM to convert this into a checklist.
Then run any document against the checklist and demand evidence. Anand's contract-analysis demo does exactly that: for each rule, a yes/no and the clause that justifies it. "Copyright and ownership permissions" → yes, section 4.1, "the author shall retain the copyright." "Quality and standards" → not found.
"Now maybe it's making a mistake; a human can go validate and also verify, 'Okay, it says this piece of text is there, is it actually there?' Very easy to narrow down."
Anand
The room fell in love with the interface — "It's such a beautiful interface." "Beautiful one." — and Debi asked the question that mattered:
Debi: "Which you just tell Claude to create. But you've obviously… knowing what we want is the most important thing."
Debi and Anand, arriving at the session's thesis
Anand: "Increasingly, yes."
Debi: "So you knew what you wanted; that's why it's exactly… which is… so knowing what you want is the most important."
Anand then undercut his own expertise beautifully:
"To be fair, the prompt was, 'Look, it should be like Excel, but fancy.' … That goes a long way."
Anand
The tools shown here: InsurLE, Policy as Code, and Straive's Contract Analysis demo (sign-in required). Anand's description of how the last one was built, when asked whether it was some special setup: "No, no, it's just Codex behind the scenes with a fancy interface. Nothing more than that."
And the anecdote that sold the whole category, from Anand R., recalling his time at IMD:
"When I was in IMD, one of our professors made a house and he said, 'I was looking at the by-laws, they were 250 pages. I read it. I'll be surprised if anyone else did.' So this would make it easier for him to do."
Anand R., IIM Alumni Singapore
"There is one really niche way of deploying applications that almost no one knows about or talks about, which is that your browser itself is a pretty powerful coding environment."
Anand, introducing the last sectionA bookmarklet is a bookmark whose URL is a small JavaScript program. Click it, and the program runs on whatever page you're looking at. The technology is ancient. Almost nobody uses it. And that, Anand argued, makes it the most under-exploited deployment target on your machine.
"You can by and large create anything where the activity is restricted to that one page."
Anand
He then opened his own bookmarks bar, which is where the session got genuinely conspiratorial.
That last one deserves its own paragraph, because it is the funniest and most honest thing said all session:
"At Straive, a lot of times the team would reach out to me and say, 'Anand, can you create a demo for X?' See, ChatGPT can already do that. 'Yeah, yeah, but if we tell our client that ChatGPT can do it, they won't buy our services. So I want you to create something that does exactly just this one little thing which I know ChatGPT can anyway do, I won't tell them.' … Now, I'm lazy. So what I do is give them a Bookmarklet."
Anand
The constraint is real, though. Anand was careful about it. Multi-page journeys — scraping second- and third-degree LinkedIn connections, or crawling every TED talk's detail page — strain the model. "Bookmarklets don't… they might work. I'm not sure if I've tried and succeeded."
And when a participant pointed out that LinkedIn lets you download your data anyway, Anand gave the rule for when a bookmarklet is worth building at all:
"The download does not have, for instance, the list of invites that I have not accepted. … And you're right, if the export is there, use the export. This is only to solve the problem that is not yet solved."
Anand
Many of Anand's single-purpose browser tools live in one public collection — mostly LLM-generated, all single-page:
tools.s-anand.net — "a collection of single page web apps, mostly LLM generated." The living evidence for "single-page HTML application."
Anand unhid a new form question mid-session — "On a website you use often, what's one tiny thing that repeatedly annoys you?" — collected half a dozen answers, and then did something better than building one: he told the room which ones were possible, and why.
| What the room asked for | Verdict | Why |
|---|---|---|
| Zoom transcript capture | Possible | If Zoom runs in the browser without launching the desktop app. "Bookmarklets work on the browser." |
| Ads I have to click away | Possible | "You could just have one button and automatically have it click on your behalf. This is a good use case." |
| Summarising WhatsApp / Telegram chats | Possible | Two steps: extract, then paste and ask. "This is in fact my workflow; I don't read group conversations at all. I convert them into podcasts and listen to them." |
| Stop me after 10 minutes on Instagram | Easiest of the lot | Debi: "But that I think you already have phone settings for that." Anand: "Probably not on the browser desktop." |
| View a LinkedIn profile without revealing myself | No | "A Bookmarklet is limited to what the browser can do. And if you're on the browser, you're logged in as yourself onto LinkedIn, then you are revealing yourself." |
| Scrape 2nd/3rd-degree LinkedIn connections | Doubtful | Multiple pages. "Multiple pages are a little difficult on something like a Bookmarklet." |
The room picked Instagram. Anand shortened it from ten minutes to twenty posts for demo reasons, and made one important adjustment mid-thought:
"Now we don't need any major coding tools or anything. Claude can do it, ChatGPT can do it… [but] we will need a coding agent so that it can test. Otherwise, it won't be able to see the browser and test it out."
Anand
"Create and test a Bookmarklet that will make sure I don't read more than 20 posts on Instagram in any five-minute duration."
…then, typed mid-run: "I have a window of Instagram open, you should use that to test."
He flagged the one thing that might not reproduce on your machine: he has browser-control extensions installed for both Claude and ChatGPT, so the agent can actually drive the browser to test its own output. Without that, a participant noted — "But it can write the code." Anand: "It can write the code, correct."
"Like with most things, the trick is knowing that there is something called a Bookmarklet. After that, the creation, the execution, and all is simply tell ChatGPT or Claude or whatever to do it; it will do the rest of it."
Anand
Click to open the full-size image
Open the full-size comic page →
The bottom strip carries the session's own summary: "Ask better questions. Give AI access. Shift formats. Code the rules. Prototype. Automate the small stuff. The star isn't the code — it's the capability you unlock."
Anand had a sixth theme queued and never got to it. "We did not cover benchmarks, that's okay, leave that aside." It's worth covering here, because it's the one that turns opinions about AI into evidence — and because the material behind it is unusually good.
It starts with a piece of prompt folklore that spread across developer Twitter. Andrew Carr had been telling models to "only report to me in ASD-STE100 Simplified Technical English" — a controlled English standard from aerospace manuals. Ben Sehl proposed making it permanent:
"Adding to every AGENTS md file for the rest of time. (h/t @richardpenner for the self-referential explanation on what ASD-STE100 is)"
Anand's reaction was not to adopt it. It was to ask whether simpler writing makes the thinking worse — and then to actually test it.
He first asked Claude to design the experiment, which produced a piece of methodology advice worth keeping:
"Pick tasks where quality shows up as content, not vocabulary — otherwise you just measure style preference."
"The ignore-style paragraph is load-bearing. Drop it and the judge will score fluency, and simple English will lose (or win) for reasons that have nothing to do with thinking quality."
"If plain language costs you something, it shows up as a named missing item — a dropped confound, a dropped exception — not as a lower score."
"Also keep an open mind on direction. Forcing plain language strips padding, and some answers get better because the weak reasoning stops hiding behind vocabulary. A tie is a real result here, not a failed experiment."
It wasn't a tie. Anand ran six tasks, with and without the suffix, each pair judged twice in both orders to cancel position bias. Of 84 judgements, all but a handful went the same way.
| Task | Sources checked, without | Sources checked, with "Answer in ASD-STE100" |
|---|---|---|
| Model benchmarking | 66 sources | 44 sources |
| Evidence and judgment | 123 sources | 84 sources |
| Adversarial system design | 97 sources | 26 sources |
"Don't simplify the writing initially. Let it think. THEN, ask for a simple explanation. … For me: I shouldn't invoke my writing and speaking skills along with other thinking skills."
Anand, Simple writing hurts thinking, 1 Aug 2026
The second piece of unused material is a twelve-stage prompt-optimisation run on something mundane — the script that writes the summary and tags for every post on Anand's blog. It is the best available answer to "how much does prompt engineering actually matter, and what does it cost?"
The whole experiment — twelve prompt generations, dozens of API calls, five diverse test posts, an independent judge, then a human judge when the independent judge proved to be measuring the wrong thing — cost about 4.3 cents.
Four findings survive being lifted out of context:
And one caution that belongs next to every "be more specific" prompt tip ever written:
"Specificity increases usefulness and searchability, but also creates more facts that can be wrong."
From the summarize.py benchmark session — read the export
The proof was a hallucination: asked for more detail about a folk puzzle, one run turned the farmer's daughter into a zamindar's daughter. The bland version had no opportunity to make that mistake. The fix was an explicit final pass — silently check actors, numbers, relationships and causal claims against the source.
There's also a nice bit of accounting hygiene in there. The script had been printing costs using stale prices, overstating every run by 5× — and the actual bill to regenerate metadata for all 3,027 posts and pages came to about $2.26. The expensive resource wasn't money. It was wall-clock time: at one worker and ~3.9 s per call, the full corpus is a 3.3-hour job.
Anand's own summary, delivered while a bookmarklet was still compiling in the background, comes in two parts.
"More importantly, one, if you know that it is possible, then a huge capability gets unlocked. A big part of therefore what we need to do is keep asking things like, 'Look, I don't even know what I don't know. Tell me what might be useful for me that you think I'm not even aware of.'"
Anand
"Secondly, the difference with regard to coding is not that code is an artifact that has value, but what you get through code has value. … What seems to be making a bigger difference is access. Can it access the context on your system, the tools on your system, the permissions on your system… And even that distinction is fading. So treat code as something that anyway is part of every request if needed; it is not a big thing anymore."
Anand, closing the four-part series
The last word, though, belonged to a participant — and it was about these write-ups:
"I find the write-ups you send later, that's very useful just to revise through the entire thing. I had actually missed two of the sessions, so for that you can almost catch up completely. … the recording takes longer to go through, whereas you read through it and it just gives you pretty much everything. I don't need two hours to listen to the whole recording."
A participant, closing the series
Which is, come to think of it, the same argument as the whole session. The artifact isn't the point. The thing you get from it is.
AI Unboxed 4 · Vibe Coding · 22 August 2026