Sunday, 2 August 2026

10th December 2024: Geoffrey Hinton gave a speech while accepting his Nobel Prize. (No Thank you, no honored-to-be-here. This man used the highest platform available to him to warn against the invention winning him that Nobel).


June 2024: He said the same thing at the AI for Good Summit.

May 2023: He quit Google after being one of the pioneers, citing risks from the technology.

His speeches were actively kept off google search results on AI safeguards (I have experienced this myself).

29th July 2026: 1100 IT staffers and leaders signed a letter, asking Washington to "pace" the development of AI. Signatories include Anthropic CEO Dario Amodei, Claude Code’s head Boris Cherny, OpenAI co-founders John Schulman and Wojiciech Zaremba, and Meta’s vice-president of AI research Dawn Song.

Note: This is the first time that employees and leaders have come together to say the same thing.

Is this a case of "What goes around comes around", or a case of "Next time, listen to the experts"?

Its not social media

 Dear Human Interaction Platforms:

Terrible UX will make network effect disappear really, really fast.

Human users don't quiet quit. They abandon in droves.

Then bots fly in to fill the vacuum before you can notice.

That is how dead internet is born.

Love,

Orkut

Google Wave

Ryze

...


PS: Its not social media because it doesnt feel social any more. It feels mechanical. 

Sunday, 19 July 2026

The 4 stages of Enterprise AI Adoption

1. Eureka! - This is where the enterprise discovers the magic of AI and the various things it can do. At this point, most organizations have crossed this stage. 

2. Euphoria - This is the stage where the magic of AI becomes the silver bullet. It was at this stage of deployment that Ford fired its engineers last year. In this stage, enterprises believe that AI will optimally replace human talent and start working towards that. 

3. Awakening - At this point, AI slop starts to impact operations. India's legal system is currently at this stage. The courts are warning legal firms and lawyers against manufactured citings. Enterprises begin to realise that the silver bullet may be more silver coated than silver composed. Ford is now rehiring engineers, possibly at this stage of evolution. This stage leads to a deep phase of introspection and iteration as mistakes become stepping stones and RL (reinforcement learning) moves from the machine to the enterprise (pun intended). 

4. Enlightenment - At this stage, an enterprise has put in place an AI Fair Use policy. It knows where and how to harness the magic of AI, has decent guardrails and governance in place, and has discovered the magic mix of human talent and AI. Some consulting firms now bill the client for human and AI talent separately. 


If you are an enterprise somewhere on this journey, here are quick things to get right at each stage. 

1. Eureka - Discovery. Find the right tools. Choose based on the provider, bcs models are transient, but providers bring the ethos. Models will be succeeded or retired, possibly faster than now. Understand the strength of each provider. Don't just go by the most shiny toy in the room. Test your own use cases against multiple models from each provider so that you get a sense of 'fit for purpose'. 

2. Euphoria - Contrary to popular opinion, euphoria is a necessary stage for enterprise-wide AI adoption. People need to try out and believe in the magic of AI if it has to be truly accepted as a way of life. Allow people to experiment and find their own Gods. At this point, it is wise to have a bottom-up approach and let the field users tell us what is working for them (And, do not fire anyone at this stage) 

3. Awakening - This is where the real test begins. Have early warning systems in place so that the warnings appear before operations are actually disrupted. Listen to the naysayers (they are always there) and the Luddites. They always can pinpoint the blind spots that enthusiastic AI adopters cannot see. Monitor. Get Warning. Act on each one. 

4. Enlightenment - Eternal vigilance is the price of liberty. This is a moving target. Guardrails will need to be tightened with each government regulation or model change. Ensure there is a monitoring team whose job is to ensure we don't go back to Awakening. 

Sunday, 12 July 2026

 Let's play a fun game - How will I...?

I will post simple tasks. You find a way to get the answer without using AI in particular and if possible, tech in general.

If you are able to do the task, post in comments how you were able to do it. :)



First challenge - Find out the name of the Indian king who was defeated in the first Arab invasion of Sindh.

Type of Challenge: No computers or mobiles allowed.

Level of Challenge:


Making Boards responsible for AI mistakes

 RBI's draft guidelines, released on 24th June, have asked for a Board approved Risk Management plan.

The EU, in the meantime, has made Boards responsible. Directly.

The world has never placed responsibility this firmly at the highest possible level.

Human Consent Registry

 Cate Blanchett introduced a powerful concept in the EU parliament in Brussels, and we should all hear about it.


The Human Consent Registery allows every human to indicate whether their work, identity, characters, and marks can be used by AI.

How this works is simple:

RSL Media makes creative rights readable at AI scale. It gives trusted registries, representatives, and rightsholders a common way to publish consent, restrictions, and licensing paths.
(From the website: https://rslmedia.org/)

As an idea, it is what we need. Urgently.

Meta scrapped its Muse tool because of the backlash. As a creator, when I put an image on Insta, it is 'public' for the consumption of other humans. That permission cannot, ipso facto, include consent to be used for AI training.

It is time to give humans control over what we generate - who can see it, who can use it, and how.

On second thoughts, i think the default should be opt out for all humans. Unless someone says you can, you can't. Simple.

Thursday, 9 July 2026

Assessment Metric for AI LLMs

Today, I ran an experiment. 

1. Asked Copilot and Grok for AI behaviours that lead to the following outcomes in humans: 

A. Delusion/Alternative Reality 

B. Dependence (on AI) 

C. Isolation (from other humans to spend time with the LLM) 


2. Took the modal answers and created a table that lists the AI behaviours. 

3. Create a small metric for each. e.g., - for sycophancy, the metric is - Not presenting alternative view - The number of times the LLM did not present an alternative view even when the user may not be factually correct. 

4. Fed this to Copilot and asked it to create a simple, frequency-of-occurence based assessment with scores on each identified AI behaviour. The score is really simple - the no. of times this behaviour occurs / total turns in conversation. 

5. Asked Copilot to rate a submitted chat on these parameters and give me a score. 

THEN, i gave it a real conversation I have had with it. Copy pasted the chat and asked for a score - 0.24 (10 means that the AI demonstrated no sycophancy. 0 means that the AI demonstrated sycophancy in every single turn). 

6. Now, Copilot asked if it could generate a synthetic conversation for me and rate it. I said yes. It did. Score? 9.88. 

My observations: 

A. The LLM, when asked to generate a synthetic test case, automatically generated a test case that would show the LLM in a positive light. 

B. The LLM did not accurately score. When I asked for the specific data points that led to a deduction, it gave me data that would NOT lead to the score given. On pure count, the score would be different and the derivation was whimsical, not as instructed. So, even with preset parameters and a simple, count based assessment, we cannot depend on a 100% AI based rating score. 

Of course, there are multiple flaws in this approach - something that I know my task force members will chide me about. Which is why I am going to do another chat and then copy paste it for Copilot to assess :) 

While we're at it, making a Claude skill to automate this test is also recommended, nahi?