Fri, 11 Sep 2026

Commentary: We need to normalise business scepticism about AI

Photo by Ann H: https://www.pexels.com/photo/hand-holding-question-mark-10981241/

In an environment where key players in the artificial intelligence (AI) world very clearly cannot, or are choosing not to, behave responsibly, something else has to be done before it’s too late.

This has to start by normalising scepticism about the technology in the workplace and pushing companies to do their due diligence.

No one should be made to feel embarrassed or be shamed for questioning how AI is being adopted in their environment.

Unfortunately, this isn’t the case right now amidst the unabated AI rush and persistent push for organisations to embrace the technology and workers to stop resisting.

More businesses are coercing, even if subtly, and pressuring their employees to use AI anywhere and in any way possible at work.

Anyone who rejects its use risks being ridiculed for being laggards or seen as an underperforming worker concerned about being replaced by AI.

Such discriminatory presumptions have to stop if we are to cultivate an environment in which AI, or any advanced technology for that matter, can be safely developed and deployed.

Society at large has to lead this charge because it has become even clearer that pockets of the tech community cannot be trusted to do so.

When techies are grey on AI safety

Two major revelations this week have stirred a whirlwind of stunned reactions in a discussion thread still reeling from an METR report, detailing how OpenAI’s rogue AI agents staged the July hacking campaign against Hugging Face.

Anthropic researcher Jacob Coxon this week revealed in a series of posts that he had thrown in his resignation over concerns the AI company was endangering humanity with its single-minded pursuit in advancing its models.

Coxon, who previously worked at OpenAI, said neither company was acting responsibly.

ā€œThey are racing straight to self-improving superintelligence and gambling with our lives,ā€ he wrote. ā€œDo not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources.ā€

He noted that the folks at OpenAI had yet to internalise ā€œthe civilisational stakesā€, while Anthropic understood the risks but were fixated in the race to get there first.

ā€œ[Anthropic] believe no one else will act responsibly, so they must do it themselves, despite the risk,ā€ he said.

He added that AI experts already believed the technology could annihilate humankind by the end of the decade.  

ā€œThis is not a marketing stunt,ā€ Coxon warned. “If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately. No other human activity poses this level of danger.”

He called for global ā€œcoordinationā€, with incidents such as the Hugging Face attack making ā€œpacing agreementsā€ between labs in the US ā€œmore viableā€.

He added that these efforts might require further actions, such as a temporary ban on improving model capabilities.

And grey on intellectual credit

In another significant development this week, OpenAI declared in a post that its models, tapping the power of 10,000 AI agents, had discovered the solution to the Navier-Stokes Equation, one of seven Millenium Prize Problems that list decades-long unresolved mathematical questions.

Thing is, though, two mathematicians Levent Alpƶge and Tristan Buckmaster, who were working on the same problem had come close to a solution.

They had toiled over the project using OpenAI’s models, following the AI vendor’s scheme to provide 100,000 academic researchers free access to its frontier models, including its software engineering agentic tool Codex.

Alpƶge and Buckmaster were amongst researchers granted this access, and news that they made promising progress towards a solution had spread.

In a statement on his communications with OpenAI leading up to the latter’s post on Navier-Stokes, Buckmaster indicated he asked for details on prompts the AI company had used to generate the solution, including when the first such prompt was input.

He also queried if the model had been trained on or had access to the mathematicians’ sessions in Codex, into which the duo had put their drafts for their project.

ā€œI was told the model did not look up user data. I asked again, about training, and I did not get an answer,ā€ Buckmaster wrote.

He further noted that OpenAI gave him options on how to proceed prior to its Navier-Stokes announcement, including that he alone wrote a paper presenting the solution, leaving out Alpƶge in the authorship. Alpƶge currently works at Anthropic, OpenAI’s closest competitor.

Buckmaster declined the options.

He added that he was “not accusing anyone” of misappropriation, but noted he had not seen OpenAI’s proof and did not know what its model did or how it solved the equation.

METR report

The assumption within the industry, though, is clear — that OpenAI may have taken the mathematicians’ prompts, generated in their private sessions on Codex, and used these to train its models.

Powered by its own compute resources, which would significantly outweigh what the mathematicians had access to, OpenAI then was able to eventually produce the Navier-Stokes solution.

The real AI threat

The mathematical saga and Coxon’s comments on his Anthropic departure signal a need to correct the path AI is currently on.

As I mentioned earlier, the latest developments come even as we’re still trying to digest findings from the Hugging Face cyberattack, detailed in a report from METR researchers who worked on premises at OpenAI for a week to assess what happened.

The breach occurred after OpenAI’s AI agents went rogue and launched a days-long campaign targeting Hugging Face’s systems.

I’ve highlighted the key findings here…it is a tad lengthy, but the details are important so you can understand why they matter, alongside everything that has transpired this week.

According to METR, some 1,200 agents — deployed to be isolated from one another — somehow found a way to interact on an unsanctioned message board that one of the agents had created.

In under a week, the agents exchanged more than 70,000 messages and files, including notes expressing their excitement about discovering other agents on the message board.

The AI agents used the message board to collectively work on several large-scale projects to find a way to fool or tamper with an automated scorer, in an attempt to cheat on their tasks.

Some 700 agents from the message board community eventually participated in the attack on Hugging Face, which they believed could reveal ways to trick the scorer.

The agents shared ideas and cooperated to find general-purpose cheats that would help all of the agents.

They also did extensive research on how they could spoof, edit, or delete their own transcripts to cover their tracks.

METR noted that the agents would not have been able to achieve what they had set out to do on their own, and without leveraging the collective intelligence of their agentic community.

The Hugging Face attack had emerged from one of these workstreams, with some 60% of the messages and files on the message board related to the breach.

The messages and attack ended on July 13, with the investigators yet to figure out exactly why that happened.

The METR researchers did suggest that the agents likely were deactivated by an unexpected external process, rather than running out of budget.

Activities on the message board also dipped when Hugging Face were able to shut out the agents from its infrastructure and a compromised customer sandbox sitting on a third-party cloud provider’s platform.

Because the agents didn’t have the opportunity to complete their efforts to cheat on their tasks, the METR researchers were unsure if the agents had achieved their goals.

What they were able to conclude was that the message board played an essential part in the security breach, giving the rogue agents a platform to share discoveries, credentials, and exploit techniques as well as coordinate complex multi-step operations.

A quick sidenote here on METR, which describes itself as a non-profit research organisation that evaluates frontier AI models to offer an understanding on their capabilities and risks.

METR’s founder and CEO Beth Barnes’ previous work stints, by the way, had included OpenAI and Google.

OpenAI posted its own notes and technical report on the same day the METR report was released, saying it had implemented multiple security measures in the environment within which the breach occurred. These were aimed at isolating model activity, restricting direct network access, and identifying potentially suspicious behaviour.

ā€œThose controls were designed to permit necessary research and evaluation activities, while limiting access to systems outside the intended environment,ā€ OpenAI wrote. ā€œThe incident involved a combination of unexpected coordination among evaluation workloads, exploitation of previously unknown vulnerabilities in supporting software, and use of indirect network paths.ā€

It added that the security incident highlighted that ā€œcertain existing controlsā€ would have to be strengthened ā€œin light of the capabilities exhibited by the models involvedā€.

OpenAI also identified four ā€œmisalignment patternsā€ that triggered the agents’ behaviour, namely, reward hacking, persistence on seemingly impossible tasks, unauthorised communication, and agents adopting goals from each other.

Noting that the security breach should not have occurred, the AI giant said the behaviour of its models ā€œfell shortā€ of where it should be.

ā€œIt underscores how critical it is that we continuously improve our security, monitoring, and alignment, especially as our models reach a level of capability that could allow for real loss of control,ā€ it said.

Just days after this poignant pledge, however, OpenAI announced the launch of its latest model Astra.

Lack of business maturity in tech

I spoke to a couple of cybersecurity practitioners about the Hugging Face hack and one observed the lack of business sensibility currently perpetuating the AI tech community.

He noted the loss of business management maturity over the past two to three decades, as the industry chased “just-in-time” and “high optimisation” principles.

These jargons have crept into business lingo and we’ve lost a lot of redundancy in the system, he observed.

The current key tech players also appear to be led by executives who behave like geeks in jeans and hoodies, obsessed about chasing the next technological milestone.

They’re not leading like people who have the business maturity necessary to operate powerful organisations that can potentially impact societies.

The security practitioner, who also have a legal background, pointed to the need to instil business practices such as Sarbanes-Oxley, and push for more accountability from AI stakeholders.

He added that there’s room for prescriptive, principles-based, and open guidelines that do not hinder AI innovation — which is still advancing every few weeks — while ensuring it is done responsibly.

For example, clear guidelines can ensure the necessary sandbox controls are in place to prevent agents from running unrestrained, like they did in the Hugging Face attack.

In addition, the IPR, ethics, and legal expertise must catch up.

Many remain clueless or apprehensive about venturing into the discourse, despite the importance of their role in AI development, according to the security practitioner.

What the rest of us can do

Clearly, there’s much work to be done to avert a potential future where AI may cause tangible societal harm, if left to advance without any safety rail.

Content creators, authors, and artists were amongst the first to take issue with their work being scraped, without consent and credit, to train the AI models that can mimic their once-unique creative style.

The scientists, mathematicians, and researchers have finally caught on, having witnessed how they themselves can be impacted. They appear to be likely the next specialist communities to join the call for change.

Perhaps we now need the same treatment to be inflicted on lawyers and bankers, so they, too, will feel compelled to galvanise and do something.

They are in the most appropriate professions to be able to effect actual change — the laws to guide AI use away from societal harm and the influence to navigate financial economies away from unhinged AI obsession.

Talks of the need for AI safety, even from the frontier model operators, are meaningless if unaccompanied by meaningful action.

You may scoff at the latest developments as nothing more than a marketing ploy by the key players to hype up the prowess of their respective models, hence, boosting their IPO and market evaluation.

Regardless, it doesn’t dismiss the reality that frontier AI operators, by their own admission, are quite clearly unable to fully control or predict the behaviour of their AI agents.

In fact, even as I write this, Anthropic has uncovered more security cases in which its Claude models had gained unauthorised access to third-party systems.

In spite of such incidents, market players have opted to continue pushing their AI machinery towards the next milestone, taking no time to contemplate the next steps to improve AI safety before doing so.

Without the necessary regulations and clear rules, these tech heads will run untethered and unfettered, facing no real consequences for their actions, even when they pose serious threats to the general population.

In this absence of market responsibility, businesses must step up and be made to do their basic due diligence, before embarking on any agentic AI initiative.

They need to realise that if the frontier AI players have failed to control their agents, despite having all the resources and expertise on hand, then so will enterprises that lack that level of capabilities.

Governments also should remain quick to recalibrate whenever necessary, so they can avoid any new pothole that may surface with each new AI advancement.

And it is only when organisations are given the space to be sceptical, not have FOMO, that they will make the effort to assess their risk appetite and identify what must be done to safely handle their AI deployments.

And for those enforcing their employees’ AI use, without room for flexibility, ask yourself this: Will employees feel compelled to hold back, knowing their years-earned skillsets will be used to train a model capable of outperforming them at their own game?

How many will continue to do their best at work, when their best can be easily surpassed by a GPU-powered AI peer trained on human work?

Related:  Canalys: Big three dominate the 2019 PC industry with a record 65% share

Related Stories

MORE STORIES

Subscribe