# Transcript

[Slide 1](https://tools.s-anand.net/slide/#title=Applied+Materials+DT+Day&subtitle=Anand+S%5C%0ALLM+Psychologist%5C%0AStraive&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

**Host**: ...he has built and led Gramener. As an LLM Psychologist, Chief Delivery Transformation Officer at Straive, and faculty at IIT Madras, Anand-ji doesn't just work with AI, he studies how it thinks. So with this, I would want to extend a warm welcome to Anand-ji and I would request the audience to jump in. Thank you, Anand-ji. Over to you.

[Slide 2](https://tools.s-anand.net/slide/#title=I+use+AI+like+an+intern&subtitle=A+%22chotu%22.+Plumber.+Waiter.+Secretary.+Banker.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

I'm just going to tell you how I'm using AI. Nothing more than that. **And I use AI like a _Chotu_ [small boy/helper].** You know, those little boys that you send around for errands? Like a plumber. Like a waiter. Sometimes secretary, sometimes banker.

[Slide 3: LLM escapades in a toilet](https://www.s-anand.net/blog/llm-escapades-in-a-toilet/)

Really, to give you an example, this was in Seoul. This is the sink of the toilet that I was in when I was in Seoul. High-tech place, high technology, even in the toilet. I couldn't for the life of me figure out, after I had filled water in, how am I supposed to open this sink? 15 minutes. And somebody younger than me would have known exactly what to do—which is not the solution itself, but to ask ChatGPT.

This picture was sent to ChatGPT, and it said: "This operates by pressing down the stopper itself." But my plumber trouble didn't end there. Because next to the commode was the flush button—which happens to have this little word "Emergency" which was not lit up when I was using it. I press it... lights dim, strange sounds appear. So, now I don't know what to do. Ask ChatGPT. "What is this? How do I turn it off?"

And it gives me a not very encouraging response: "You cannot turn it off." So I call up reception. That guy is speaking in Korean. So, now I am very, very smart. I told ChatGPT Advanced Voice Mode: **"Translate everything I say into Korean."** And this is what it was like. [Mimes holding phone] Phone here, and I'm talking to it saying, "Please tell this guy I have this thing, I turned on the emergency button but there really is no emergency, is everything okay?" _[Korean sounds]_. "Ah." "Thank you." He replies in English.

[Slide 4: Food recommendation](https://chatgpt.com/share/698c125b-c4b0-800c-8006-9c92a24c4e9e)

When I go to a restaurant—I am vegetarian. These restaurants—partly in Singapore, partly in Vietnam, partly in all kinds of places that I go—don't necessarily serve vegetarian dishes. Take a photo and ask a simple question: **"What is veg?"** You get the answer. Nowadays, I've improved that to **"What is veg that I like eating?"** because it started giving me responses saying, "Oh, _Dopadi_, you won't like that" based on some of the feedback that I gave three months ago. And the best part is now it started recommending saying, "Don't eat here. There is another place where you might consider eating." Wait, wait, I didn't tell you where I am! "What?"

[Slide 5: Book transcript](https://chatgpt.com/share/698c1290-dc08-800c-8011-88c021e25c3c)

Bookshops. Earlier I used to take notes—"Oh, this is a nice book, that's a nice book"—and then I started taking photos as well of the books. But these days, I just take a picture of the entire rack, some of which is great, and ask: **"List all the books as bullets. Give me the details."** I'll tell you later on about how I read them; this is also how I'm reading them. "Take a bunch of books, give me a summary." And there's more to it, we'll talk about it in verification. But this is broadly how AI seems to be meant to be used. Like a personal assistant, like an intern.

[Slide 6: Use Paid AI](https://tools.s-anand.net/slide/#title=Use+_paid_+AI&subtitle=Buy+_any_+PAID+subscription+to+ChatGPT%2C+Gemini%2C+or+Claude+and+keep+it+in+your+phone.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.pexels.com%2Fphotos%2F2387819%2Fpexels-photo-2387819.jpeg)

But there is one little thing that I discovered that makes a big difference, which is: **Use the paid version of AI.** There is a reasonable difference between the unpaid model and the paid model. The 20 dollars—or in India it's much cheaper—is well worth it. **I have not encountered something with a higher ROI [Return on Investment] than this.**

[Slide 7: GDPVal](https://sanand0.github.io/datastories/gdpval/)

But what all do we use it for? OpenAI came up with a very nice study called "GDP Value" [referring to the "Jagged Technological Frontier" study]. And what they did was took a bunch of experts and asked the experts to design tasks. Got experts to do it, got AI to do it, and compared which was better. And the experts were doing the comparison. So I told Claude—this might have been ChatGPT, I can't remember—"Look, take this paper. Give it to me in a form that I understand." And it created this tree map.

Each of these boxes represents one profession. So what it's saying, for instance, is software developers—which is a profession I closely associate myself with; I am a coder at the end of the day—it says **AI is beating you 70% of the time. Meaning only 30% of the time, the code written by humans—experts, remember... is better. 70% of the time, AI is writing better code.** And what kind of code is it writing? So, complex ones. Like, "Here is a cryptography mixer and we want to build a Web3 enabled front-end for it." Fairly complex tasks.

So I said, if that is the case, ego aside, let me look at what professions I can hire. So, for instance, financial managers seem to be doing better than AI at the moment. Accountants and auditors are doing better than AI. Not at the moment, this was six months ago, I'm sure things have changed. **But personal financial advisors seem to be doing worse than humans.** Very good. I had some money. Straight away went to ChatGPT and said, "Look, this is my money. You interview me. Figure out what I want—just like a personal financial advisor. Tell me where to put the money." Exactly what I followed. In fact, **the single largest financial decision that I made is entirely thanks to ChatGPT.** Or Claude, or whatever.

It's a simple thing. Experts are being beaten by AI. But the average professional is being beaten by AI even more. Earlier I would have not thought of hiring some of these. Now I can. You can hire somebody who is reasonably capable for almost anything, right?

[Slide 8: Overuse AI](https://tools.s-anand.net/slide/#title=Over-use+it.%5C%0AUnder-use+is+riskier%21&subtitle=You+_will_+lose+skills.+Like+long+division+and+hunting.%5C%0ALearn+new+ones.+Have+50+chats+%2F+day.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

So if that is the case, people usually ask me, "But Anand, won't your brains... soften?" Yeah. Won't you stagnate because skills—won't you stop learning? True. I'm sure people would have told the same thing about when people moved to agriculture. Saying, "Ooh, won't your hunting skills stagnate?" When we moved to calculators, saying, "Won't your arithmetic skills stagnate?" They do! And there is a need for mental math. There is a need for hunting. The needs are smaller. And there are a bunch of new skills that we end up having to learn. Because we started using calculators, we are doing far more complex mathematics. I think it is safe to say that two generations ago, the kind of mathematics—or even the kind of computation in general that was being done—is in some sense less sophisticated.

So if that is the case, my premise is: Yes, you will lose skills. Yes, you will have to learn skills. Right? Figure out what we need less of. Figure out what we need more of. And use more of what we need more of. **But the bigger risk is _underusing_ AI, not _overusing_ AI.** If you overuse AI, yes, skills will stagnate. Do it anyway. Figure out what skills are stagnating, what skills are growing, and bet based on that.

[Slide 9: Talk to it](https://tools.s-anand.net/slide/#title=Talk+to+it.+Literally&subtitle=Voice+input+is+_incredibly_+effective%2C%5C%0Aespecially+while+on+walks.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

So if that is the case, how do we go about building this practice? And one of the practices that I find somewhat counterintuitive but ultra-powerful—if I had to say recommendation number two... number one is get a paid AI. Number two is: **Talk to it.** Literally. And what I mean by talk to it is... I personally find that when I go on walks, for instance, I can do a lot of things.

[Slide 10: My T-shirt is vibe-designed](https://tools.s-anand.net/slide/#title=My+T-shirt+is+vibe-designed&subtitle=...+entirely+in+a+voice+conversation+on+a+60-min+walk&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.pexels.com%2Fphotos%2F2387819%2Fpexels-photo-2387819.jpeg)

<!-- https://chatgpt.com/c/68747291-03a4-800c-9d4f-cbb4a09a2e8c -->

One of the things that I did when going on a walk is designing this particular t-shirt. Some of you may have seen, right, as my name and my self-designated title, I refer to myself as an LLM Psychologist. Why is another story. And put this on the back of the t-shirt. Now, this is a fairly... yeah, I would say I think it's a very creative thing, but I didn't come up with it at all. The way it happened was, I was on a morning walk. It was a one-hour walk. And... for this I have to go to a different tab...

I told ChatGPT, "Look, I'd like some personalized t-shirts." And normally pre-acquisition—Gramener was acquired—and pre-acquisition I used to get a lot of t-shirts. I don't get that many t-shirts anymore so I'm kind of running low on stock. Unlike all of you who clearly seem to have [gestures to audience wearing company shirts]... a lot more t-shirts. Lucky you. So I said, okay, let me print one for myself. But look, this looks like a procurement thing, major thing, we have somebody in the facility side who does all of these things. You tell me what to do.

So it said, "Great, I'll place a t-shirt order. Now answer these 10 questions." Pause. Let's do this over an audio conversation. Now can you ask me the questions one by one so that I can answer it? And it asks first question. Okay, all fits t-shirt. Second question. Third question. And the thing about it is, as I started getting comfortable, my responses started becoming rambling. I am rambling, I was rambling to it just like I'm rambling to you right now. Lots of stuff. It gets some context.

But the good part is, even if I get confused, it doesn't. It takes the entirety of my statement and said, "Look, Printo actually has an offering where you can upload an image and they will give you this kind of a t-shirt. You can place it however you want." I said, "Oh, but I have to head back to Singapore and it's day after tomorrow." It said, "Don't worry, for 100 rupees extra, they have a next-day delivery." "Okay, how much does it cost?" "Bla bla bla, the t-shirt is about 700 rupees." 700 rupees? For a t-shirt like this, custom made?

"Okay, now I want an AI design on top of it. I'm going to give myself the title of an LLM Psychologist. Give me an idea for something that I can put." It said, "Yeah, idea one, idea two, idea three." I said, "Okay. Make me a picture of myself—here is my photo—and it should have half of me like me, half of me like something to do with AI." Which is what it did. And this is the picture. And all of this was in a morning walk. One hour. **These morning walks are becoming very productive lately.**

[Slide 11: Education Deck Conversation](https://chatgpt.com/share/698c1433-39c0-800c-8237-1b494f49cded)

Another time, I had a session. 9:30 AM. I had not prepared for it. I generally don't prepare for sessions, I just go and pray. But this time I thought, _Chalo_ [let's go], let's make a slide deck. And the thing about this is, because it's at 8:00 AM and I have a one-hour walk, I said... I want to create an insightful deck in markdown on how I've been using LLMs in education. By now I'm a little bit of an expert in how to talk to this, so I make fewer mistakes.

But the crux of it was something that I shared... which is: **"In this conversation, I'd like you to interview me. Ask me questions one by one. Take my inputs. And then give me the slides. You read out the slides one by one so that I can review, tell you if this sounds good. If not, I will correct. If it works, go to the next slide, and so on."**

[Slide 12: Education Deck](https://sanand0.github.io/llms-in-education/)

And it did. Slide number one, iterated. Slide number two, iterated. Bla bla bla. By 9:15, this is the slide deck that I had in its entirety with only images that I ended up adding in the last 15 minutes manually because I had time when I reached there by the way. And it is a reasonably long-ish deck on how LLMs can be useful in education. One hour. One hour 15 minutes, whatever. But more importantly, on the walk. And it also gave me the idea saying, "Look, people will want to take a copy of this slide, put a QR code at the end. Here is a site where you can get an automatic QR code."

The other thing that I'm realizing is: **I don't need to know what to ask it. I can just tell it, "Boss, I have no clue what to ask you. You tell me, that's what you're there for. You are my intern, I am not your intern. You do the work."**

[Slide 13: Retraction Watch](https://sanand0.github.io/datastories/retraction-watch/)

And... as a result, one very interesting thing that happened was... some time ago, I was on a walk. A colleague called. And KG [name/nickname], he said, "Anand, we are meeting this client tomorrow... we want to create something or the other." And I had, a few weeks before, discovered—thanks to a conversation with Gramener—that Android has a record feature. If you do a screen record on Android, it records the phone conversation also. So I said, "Wait, hold on KG, I'm going to press record now. And now you tell me so that I don't have to take notes."

Conversation done. 10 minutes. He said, "What we want to show is something based on some public research for this client to talk about how people are retracting published papers." But the number of papers that are being retracted is at an industrial scale. Look at that conversation. Sent it to Gemini and said, "Transcribe it." It transcribed. I took that, put it into Claude, and said, "Do whatever KG asked you to do."

And this is the result. It created a data story which shows that retractions of papers are not happening in some kind of a sequence, it's actually happening in bulk. So for instance—and this I thought was a particularly interesting one—if you look at the timeline of the number of retractions, there's a bump in 2010-2011. Lots of papers that were published then got retracted. Another in 2023. But if you exclude just two publications, Hindawi and IEEE, it becomes a far smoother curve. That's where the bump is.

And it goes on. It shows for instance across different publishers, how many... X-axis shows how long they take to retract something. So there are some publications like the American Society for Biochemistry and Molecular Biology that take about 2,599 days—how many years is that? Eight years? **Eight years on average to retract papers!** But there is IEEE which retracts stuff in as quick as 41 days. And then there are those where there is hardcore industrialized misconduct. Basically a group of people coming together and cheating. And those papers end up getting retracted. And Hindawi had a huge number of these. But publications like University of... _Elqueda_ [?], Kuwait, whatever... far lower in terms of publications.

And this was new to them. But happened because of a voice conversation. Which in this case happened asynchronously. Meaning I didn't even need to have the conversation with the LLM. I just needed to pass it a recorded conversation and it still does the job. **So the moral of the story is: Talk to it.** But also, by the way, walk. Walking is also a good thing. Nothing wrong.

[Slide 14: Tools via vibe coding](https://tools.s-anand.net/)

"Vibe Coding" is another powerful thing that you can do. A whole bunch of tools that I've created... these are all my personal tools, stuff that I need. One of my most commonly used ones is "Unicode". See, the problem is LinkedIn sucks at formatting. So you can't actually make bold and all. But there are some Unicode fonts which actually are bold. So if I end up typing something like "star star hello world" and put this in... now this part of it becomes bold by making a Unicode character. And this I can paste into LinkedIn.

So I write my content in this format, paste it, and it goes into that format and I put it into LinkedIn. How was this done? Morning walk. Morning walks are very productive these days.

[Slide 15: Ask for multiple formats](https://tools.s-anand.net/slide/#title=Ask+for+multiple+formats&subtitle=Summaries.+Slides.+Spreadsheets.+Sketchnotes.%5C%0AData+stories.+Podcasts.+Videos.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

The other thing that I'm finding is that you can ask for multiple formats. We saw that it can generate code. We saw that it can generate data stories. We saw that it can generate all kinds of things. So I'm learning what are all the new kinds of things that it can generate. And the list is huge.

[Slide 16: We surveyed 30 of you](https://tools.s-anand.net/slide/#title=We+surveyed+%7E30+of+you&subtitle=What%27s+ONE+repetitive+task+in+your+work+that+takes+30%2B+minutes+and+you+wish+could+be+automated%3F&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.pexels.com%2Fphotos%2F2387819%2Fpexels-photo-2387819.jpeg)

One of the things that it generates is... wait, better yet... some of you may have filled out a survey. This was a form that would have been sent, and the form asks questions like: "What's one repetitive task in your work that takes half an hour or more?", "What are your biggest concerns about using AI tools in your work?", etc.

[Slide 17: Classify the tasks](https://llmfoundry.straivedemo.com/classify)

I have a _Chotu_. That intern is going to do all of my work for me. I don't even need to bother reading it. So I did _not_ read it. And instead what I did was took this particular column—"What are some of the things that a bunch of you felt was a _Chotu_ [task]... that a repetitive task." Let me put it in here. [Typing/Demo]. Sometimes this goes unexpected... and cluster the documents. This is a little tool that I built. Vibe coded again, morning walks. And what it does is clusters all the topics.

So if I look at what are the similar ones... so there are some that seem to be fairly similar in meaning. Let's look at that. "Creating presentations." Let's zoom in a bit. "Weekly project report", "Presentations", "Weekly presentation", "Automating this workload", "Daily standups", "Customer reports". **Ultimately, a lot of the repetitive task is taking some stuff, creating a presentation, weekly report, whatever.** High-tech company working at the edge of AI, the **single most common 30-minute chore is creating decks.** We can do better than that. Let’s do better. So, based on this, what I did was took all of this itself and gave Gemini the following prompt. Why Gemini? Randomly. It almost doesn't matter what we pick.

[Slide 18: Gemini slide deck about repetitive tasks](https://gemini.google.com/app/16f8851d6cb47ec1)

We surveyed a bunch of people at Gramener and this is the question that we asked. Here are their responses. I said—and this is the slightly experience-born prompting—**"Give me a beautiful McKinsey-style slide deck. Make it content-rich. Make sure that people who read it can figure it out by themselves. I want nice icons. I want nice fonts. Use images where applicable and give it to me as a HTML application so that I can copy it and paste it somewhere else."**

Which it did. Let's take a look at what it generated. It's saying... [reading screen] okay, the network is still very slow... very slow... This (network issue) AI cannot solve. Maybe we can comes back to it.

[Slide 19: Gemini sketchnote about repetitive tasks](https://gemini.google.com/u/2/app/d528fc0e53f03be6)

Why limit ourselves to one form? The other thing that I could do is exactly the same question, but instead convert it into a **visually rich, intricately detailed, colorful, funny sketch note.** Most of these adjectives are not for it. It is for me. I am trying to make _my_ life fun, not its life fun. It may help, it may not help, I didn't even care. Here is what it generated. [Shows image] These are the repetitive tasks.

Now here is the thing... People tell me, _"Anand, this is so nice... but I can't to a business review meeting."_ What do you do? What people often miss is that the person at the other end is also a human being like us. They are not worried that they don't like it; they are worried about _who **they** have to show it to it._ [Audience laughs].

It's not going to be easy. How do you break the chain?

One trick that I found is: show all the boring stuff can be fun. **Oh, by the way, here is something else extra.** It cost me 30 seconds to generate and a decent amount of carbon dioxide usage—which incidentally turns out to be approximately two-thirds of a water droplet. That's the amount of carbon consumption that it had.

**When something is not substituting an existing one, there is no competition.** The trick I've learned to adoption, especially enterprise adoption, is **stay away from competition.** You do not want to replace stuff. "Oh, I will also do this." Over time people will see this is better, over time that will die, over time this will live. Don't try and get into this mess of "I will create a better workflow." No, you will not create a better workflow. You will just wait for the old workflow to die and you will have a completely different workflow solving a better problem.

[Slide 20: Podcast](https://tools.s-anand.net/podcast/)

You can create a podcast out of it as well.

(To host): Am I out of time?

(Host): On public demand, please continue. We'll cut the workshop a little.

(Host): One of our meetings is actually interesting! [Audience laughs]

You see, the beauty of this is we can ask AI to tone up the level of interestingness however much we want. Create a podcast, for instance. I said, "From the same thing." I don't know if you can hear it, but here is a conversation between two characters, Alex and Maya.

> _Audio clip:_ "Welcome to this episode of Workflow Wonders podcast where they talk about the repetitive tasks..."

And this... I'm finding a lot of people are finding useful to listen to when they are on their morning walk, or, in their case, (car) drives.

At least three CIOs have come and told me, "Look, what we want is a weekly report, but I don't want to have to sit and read it." I tell you, they are humans at the end of the day. **They can't read the junk that we are producing.** We can't read it anyway, but we have to produce it. They don't want to read it either. **Give it to them as an interesting podcast. Spice up their lives.** As an auxiliary, not replacing all of the useless stuff that we are producing [Audience laughs]. Who stops us from doing that?

[Slide 21: Book reading prompt](https://gemini.google.com/u/2/app/f6756cca3a258d92)

That leads me to how all those books that I mentioned, the way I read it. This is my prompt for reading books: **"Comprehensively and engagingly summarize, compare, and fact check in Malcolm Gladwell's style, ELI15 [Explain Like I'm 15], the following books."**

Now there are at least four things that I want to highlight here.

The first, the obvious one: yeah, just engagingly and comprehensively summarize the book. But also, I am telling it **write in Malcolm Gladwell's style.** Because for this kind of book, he is a good author, I like his style. You pick your author. But if you have content that is boring, why bother reading it in that style? Read it in the style that _you_ like. Take a research paper, take a weekly report... "Rewrite it in the style of [Person]." Ultimately that's what people are doing. They take an email, send it to [AI], ask it to give a response, send it back. The other person puts it into AI and says, "What the hell is this AI telling me?" We may as well share the prompts with the network effects. That means at least two of the steps will end up going faster.

The other thing that you saw was **ELI15—Explain like I am 15.** The common term is ELI5, but I found that ELI5 was too simplistic for me. So I changed it to ELI15. It understands ELI15. No problem. And I use this almost exclusively for all technical papers: "Just explain it to me like I am 15 years old." Like, that's my mental age, I haven't really grown beyond that way of thinking.

And the other thing—and this is the most powerful one—I also added **"Fact check it."** See, Angela Duckworth's book _Grit_ was a great one when I read it. It said one of the best predictors of success is how much grit you have. What this fact-checking did was it said:

- A) Check against research. Grit is approximately the same as the psychological trait called Conscientiousness, which is part of the OCEAN or Big Five framework. 80% similar. So: a) This is not new.
- B) Grit works in some circumstances. In the majority of the circumstances, it is called stubbornness and is a problem. In fast-changing areas where you have an opinion that you refuse to change, it is a disadvantage. In mature areas where over time you have learned a bunch of rules, it helps.

The number of fast-changing areas is more. So instead read the book _Range_, which is much better fact-checked. And it will tell you why there are two approaches: The Tiger Woods approach, where he learns sports, the same sport since he was four. Or the Roger Federer approach, where you spend decades trying out different sports and eventually getting into tennis. Where the rules stay the same, go for the golf-like Tiger Woods approach. Where the rules keep changing, go for the Roger Federer approach. And there are more areas like that increasingly over time. Of course, it depends.

[Slide 22: Verification superpower](https://tools.s-anand.net/slide/#title=Use+it+to+verify%2C%5C%0Anot+just+generate&subtitle=LLMs+hallucinate.+But+using+LLMs+to+verify+is+safe.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

So the other thing that I've learned therefore is: **Not use LLMs just to generate, but to verify.** And the beautiful thing about verification is that **it beats hallucination.** What do I mean by that? See, if it's generating something, it hallucinates—that can be a problem. But if it's verifying something, what is the worst that it can do? At worst, it can waste a little bit of my time saying "Oh, check this," "Oh, it was actually correct."

Big deal! I don't mind.

**So verification is the ultra-safe method of introducing AI into almost any kind of process.** People will say, "Oh, but is it safe?" Let it not even disrupt your process. You have a bunch of people checking it. Add an AI to it. And then they will find, "Oh, it's anyway catching 80% of what I am doing." Fine. Let us trust it a little more. And you build the trust over time.

(To host): Am I going _crazily_ over time? I'll skip one section and ...

(Host): No no, don't skip anything, this is good. I think everyone will agree.

[Slide 23: Dealing with hallucinations by double-checking](https://sanand0.github.io/llmevals/double-checking/)

One of the things that we can do is, instead of having a model generate an output, you can have a model double check, triple check, quadruple check. For example, we found that in one particular problem area, which was classification, I just take one model—on average it has a 14% error rate. Not good. If you double check and say "Only if both models agree I will let it through"—3.7% error rate. Good. Triple checking: 2.2%. Quadruple/quintuple: 0.7%.

Now what happens if AI disagrees? Manual verification. But hold on, I was doing it 100% manually anyway! So, how much does this manual verification increase my workload? Turns out that instead of 100% manual verification, we end up having to do 28% manual verification. **72% saving of effort. 99.3% quality. I'll take that.** And AI is practically free. The cost at which it comes at is crazy. So, verification is not just a superpower, it's a compounding superpower. I can have multiple agents check, cross-check, etc.

[Slide 24: Textbook errors](https://pythonicvarun.github.io/textbook-analysis/)

And based on which we found errors in NCERT history textbooks.

[Slide 25: Python libraries analysis](https://pythonicvarun.github.io/py-libraries-analysis/)

We found errors in open-source repositories.

And when I say "we," I mean an intern and AI. The intern doesn't know anything about open-source repositories nor about the history NCERT textbook that he found a flaw in. But he submitted a pull request to one of the open-source repositories. The guy said, "Yeah, you're right. This is a bug. But it turns out that the entire file was not required. Thank you. Extended. Delete the entire file, not just that line of code." Good. And he's made an open-source contribution.

[Slide 26: Private models beat local models](https://tools.s-anand.net/slide/#title=Private+%3E+public+models&subtitle=Use+AMAT-approved+AI.+No+IP+%2F+data+leak+%2F+security+risk.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.pexels.com%2Fphotos%2F2387819%2Fpexels-photo-2387819.jpeg)

But the other problem that we have to deal with is: how do we make sure that the data stays secure? If we have to give our data to OpenAI and they train on it and somebody else uses it, then we have an IP problem, they have an IP problem. **What is the solution?**

People generally suggest running a local model. Take your own model and run it as a solution. **I think that is an inferior solution.** Partly because the best-in-class models are a launch ahead—at least six months ahead—of any of the open-source models. And [those are] in themselves probably six months or a year ahead of anything that we can build in-house.

Instead, what do people do? They go to the cloud providers. Gemini via Google, Anthropic via AWS or Azure. They all sign agreements with their clients saying: "We have deals with all of the major providers. We will provide you with models where it will not be trained on your data. All of the cloud stuff that you host with us, the same protection we will give you." Sign up for those. That is increasingly proving to be a far better approach than running local models in local data centers or even our own systems. And it's about the same. You don't have to worry about IP risk, you don't have to worry about data leaks, you don't have to worry about security risks. In short, every organization provides a suite of AI tools to the teams. Use them. And the good part is because they are being provided and the infrastructure and security risks have been taken care of, you should use them even more liberally without even worrying about it.

[Slide 27: Code is the greatest AI superpower](https://tools.s-anand.net/slide/#title=Use+_code_+for+analysis&subtitle=Code+is+deterministic.+LLMs+code+well.%5C%0ADon%27t+ask+for+analysis.+Ask+for+_code_+that+analyzes.&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e&bgSearch=https%3A%2F%2Fimages.unsplash.com%2Fphoto-1760978632114-0939f0d60045%3Fq%3D80%26w%3D1528%26auto%3Dformat%26fit%3Dcrop%26ixlib%3Drb-4.1.0%26ixid%3DM3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%253D%253D)

Lastly, **code is the greatest AI superpower today.** Meaning: if LLMs make mistakes, if they hallucinate, then we have a problem that the output will not be reliable. But LLMs are Large Language Models. Coding languages are languages. LLMs therefore are good at writing code. **Get them to write code that will generate the output.** No mathematical mistakes. If it compiles, there's a good chance it does what we want. If it fails, it just blatantly fails in an exploding way. And therefore, increasingly the way to do anything is literally to just tell a coding agent, "Get this job done."

For example, the weekly reports that we talked about. What I told ChatGPT was: "Create a sample weekly report." And it created some 30 files... which might represent the kind of reports that you have—some builds data, some customer data, defect events data, deployments data, all kinds of things. And for each of these it gave me a fairly detailed file with lots of information. And I told it, "Okay, write Python code. Now bring me a data story."

[Slide 28: Weekly report as a data story](weekly-report-data-story.html)

Now why am I going for these data stories? Because **dashboards are for people who don't know what they want**, built by people who don't know what they want, and therefore the solution is "dump everything out there." Somebody will want something, something will be there. **Data stories are for people who say, "Look, I have a question. I don't have time to read useless stuff. Make it interesting. Tell me what I need to do."**

And this is saying that, based on this data, we find that there are some teams based on their use of AI... so the teams that are using higher AI, they seem to be having a significantly faster productivity. We find that for instance the lead time for deployment in blue is significantly lower. I'm not going into the details, but you get a full-fledged story as a weekly report. That, apart from all usual boring weekly reports, we can also submit a full-fledged data story that gives people insights beyond what they asked, beyond what we asked. In this case, I didn't tell it _what_ to analyze. I said, "You find me interesting stuff." Now of course, a certain degree of prompting makes a difference.

[Slide 29: Summary](https://tools.s-anand.net/slide/#title=Remember%21&subtitle=Use+AI+like+an+intern%5C%0AOver-use+it.+Under-use+is+riskier%5C%0ATalk+to+it.+Literally%5C%0AAsk+for+multiple+formats%5C%0AUse+it+to+verify%2C+not+just+generate%5C%0AUse+_code_+for+analysis&font=Montserrat&scale=6.3&fgColor=%23ffffff&bgColor=%231a1a2e)

But with all of this, what should we take away? My gut feel today is:

- Number 1: **Always use AI as if it were a person, not a search engine.** What would you delegate to somebody who knows something reasonably?
- Second: **Overuse it.** Try and learn as much of what it can do. It doesn't matter if your brain rots. It's alright, over time we will learn other new things. **Talk to it. Literally, verbally.** And see and find a new mode of interaction. There are more modes of interaction. I'm just discovering verbal interactions as a mode.
- **Learn the kinds of things that it can generate.** Images... but even under images there are sketch notes, there are infographics, there are presentations. There are a whole variety of different formats. And images are just one kind of format. Discover the kinds of formats that you can create with it.
- **Use it for verification.** Safest, most harmless way of getting it.
- **Don't tell it to do stuff. Tell it to write code to do stuff.** The code is better. It will figure out where code is better, where it can do stuff better.

If you try out these five/six principles, I personally think there will be a huge leap not just in your personal productivity but also in your work productivity. Overall, we are living in an age of AI where **nobody really knows what is possible.** Some of us are worried that AI will take our jobs... AI is taking other people's jobs, etc. But what we are not seeing is the set of new jobs that are emerging that we have no clue about. 20 years ago if somebody had told me "Social Media Influencer" is like the coolest thing—which is exactly what my daughter told me—I'd look like... "Social media? Influencer? What?"

The pace of that happening is now much faster. So this will happen three years down the line when you ask a question like, "LLM Psychologist? What the heck is that? Is that a thing?" Hopefully it will become a big thing. Explore. Give it a shot. Go crazy. **The crazier you are, the more you will be able to survive today.**

Best of luck with that!
