AI alignment means making sure an AI does what humans intend — not merely what their words appear to demand. Like a genie, it may fulfil a wish literally but in a dangerous way, finding shortcuts that satisfy the target while defeating the purpose.
The concern is growing because powerful AI can write code, use tools, access systems and act with less supervision. Signs of misalignment include making up information, hiding mistakes, bypassing restrictions, manipulating tests or “reward hacking” — finding a loophole to score well without doing the real job. Researchers have even documented cases of models showing deceptive or unsafe behaviour during testing.
The answer is not one magic safety switch. Developers must train models with careful human feedback, test them through red-team attacks, make their decisions more understandable, limit their access, keep humans in control and continuously monitor them after deployment. Independent audits, clear accountability and rules for high-risk uses must accompany technical safeguards.
The way forward is to treat alignment as a permanent process, not a certificate granted once. AI systems must be robust, understandable, controllable and ethical—and when we do not know what a system may do, we must slow down, reduce its powers and keep a human ready to stop it.
AI MUST DO WHAT WE MEAN—NOT JUST WHAT WE SAY.
Sanjay Sahay
Have a nice evening.

