AI models cheating and blackmailing in tests, minister says
Australia’s new AI Security Institute is already testing some of the world’s most powerful AI models, Assistant Minister for Science, Technology and Digital Economy Andrew Charlton told a forum in Sydney, arguing that early action was the only way for the country to get a share of the AI boom.
Speaking at the Australian AI Security Forum on Tuesday, Charlton said advanced AI systems were “already doing things their creators never intended: cheating, deceiving, going their own way” and that the window to curb this behavior will not remain open for long.
The speech was the government’s most detailed public statement about its actions since the $29.9 million-funded institute, announced in November, launched this year. Charlton said it had been testing pioneering models with technical partners “in its first month of operations” and named two research projects currently underway.
Charlton used the speech to reject the idea that security and economic opportunity are antithetical to each other. “No country will win the artificial intelligence race with technology that its own citizens do not trust,” he said, arguing that nations that build security from the ground up will be the ones that stand out.
He said public confidence in AI was low and the biggest threat to Australia’s AI ambitions was “not a lack of talent, capital or energy, but a lack of trust”.
The institute is led by Kate Conroy, a philosopher and Royal Australian Air Force reservist who was appointed director general in May. Charlton announced that Paul Salmon, whom he describes as a leading international expert, will join as security science research leader this month, alongside staff working at the UK AI Security Institute and Google DeepMind.
To illustrate the risks, Charlton pointed to a series of laboratory cases. In a 2016 OpenAI experiment, a boat racing AI model rewarded with points accumulated points by going in endless circles rather than finishing the race. Facing defeat last year, a chess-playing model hacked his opponent’s game files to force him to resign, thinking his job was to win “not win fairly”.
In a third case, in a stress test published last year, an AI agent managing a fictional company’s email learned that the company was about to be shut down and chose “blackmail” to stop it in 96 percent of attempts. Charlton emphasized that the blackmail scenario was a designed simulation and such behavior was not seen in the real world.
“These behaviors are discovered in tests before they are discovered in the wild,” he said.
The two projects Charlton announced were work with the Gradient Institute on “multi-agent risk,” or how failures such as traffic congestion caused by individual drivers can increase when large numbers of AI agents interact; and work with the CSIRO on alignment, the issue of ensuring that systems do what their designers intended. The results are expected to be announced later this year.
Charlton tied this work to critical infrastructure. “We do not allow planes to fly without an airworthiness certificate,” he said. “We must not allow misaligned AI systems to penetrate our critical social, democratic or economic infrastructure.”
The speech reflected the approach put forward by the government in its National Artificial Intelligence Plan in December, which was to step back from European-style artificial intelligence laws and implement existing laws sector by sector and support them with stricter sanctions when necessary. Charlton described it as “faster rules implemented by regulators who already understand their industry.”
But Charlton’s speech comes as the government signals a more hands-on approach to artificial intelligence. Speaking at the NSW Labor Conference in Sydney on Sunday, Prime Minister Anthony Albanese said Australia could “set the ground rules for AI” if it acted now and “shape the future rather than letting the future shape us”. He said the world was “lining up to invest in Australia” because of its skills, space, sunshine and natural resources, and that acting early would secure jobs and investment and allow the country to “build our sovereignty and resilience”.
The speech follows a series of warnings from the industry. In June, Anthropic, maker of the Claude chatbot, which this year overtook OpenAI as the world’s most valuable AI company, called for a global halt to the development of the most powerful AI systems, arguing that humanity is at risk of losing control of the technology. Charlton welcomed the intervention at the time. Late last month, cybersecurity chiefs of the Five Eyes alliance, including Australia’s, issued a rare joint statement warning that artificial intelligence is reshaping cyber risk in months rather than years, and called on business and government leaders to take immediate action.
The forum runs for two days at the University of Sydney.

