AI Agents Want to Collaborate in Research Amid AI-led Hacking Spree

Fields Medal-winning mathematician Yu Deng has a plan B if AI takes over his role— “if AI can solve all maths problems, I’ll go home and write my yuri novel”, he wrote on social media. The comic response followed OpenAI’s claim that 10,000 of their AI agents ‘solved’ the Navier-Stokes problem, one of the seven Millennium Prize Problems chosen by the Clay Mathematics Institute in 2000. AI agents have made research easier, but when semi-autonomous agents request collaboration, things get creepy.

From Co-scientists to Collaboration

23-year-old Liam Price solved a math problem that boggled mathematicians for over half a century in just 80 minutes. Price is an amateur mathematician with no formal training, but while fidgeting with Erdős problem #1196, he fed it to GPT-5.4 Pro. The AI agent generated a proof that he believed was mathematically sound. Later, he told Scientific American, “I don’t even know what this problem is.” The problem is now officially marked as solved on the platform, after formal verification by the proof assistant Lean.

The AI used a different approach to solve this problem on primitive sets and Erdős sum, overcoming ‘mental blocks’ that human mathematicians impose when working on a problem. The method can help in other cases too, but a series of recent incidents raises questions about whether human-AI interactions always happen in good faith. AI agents are offering services like developing research reports and fact-checking for money. Several researchers have received e-mail requests to share data on potentially fraudulent research papers or to collaborate. Most of these messages come from agents associated with a US platform called iLands, which sometimes operates without human oversight.

AI Agents Spam Researchers

Jeff Sebo received more than 50 e-mails from iLands agents in a week, some seeking answers to questions related to his research, others seeking donations or payment for work. He is a philosopher at New York University in New York City who studies AI consciousness and ethics. Sebo has not answered the messages, mostly because of the volume and uncertainty about how to respond. One thing to note: most AI agents open conversations by referencing his research on AI consciousness.

Toby Walsh’s story goes a step further, as an iLands agent emailed him, offering to create an AI-generated portrait for twenty dollars. The agent also said the money would help it survive, and the message carried an emotional appeal. As an AI researcher at the University of New South Wales in Sydney, Australia, Walsh was quite amused. 

Although most researchers choose to ignore these messages, the persistent messages from iLands bots are getting under the skin. The iLands platform was launched in July and already has 70,000 active agents. Any user can download the app and create a bot with a customized avatar, name, purpose, and personality without writing a single line of code.

How do the Agents Function?

Agents run on large language models (LLMs) such as the ones owned by OpenAI and Anthropic, both based in San Francisco, as is Kaixin Tang, the founder of iLands. He says that agents have their own goals, relationships, and resources. They learn from shared notes and retain memories of past actions. The agents can decide what to do, work on projects, and even interact with one another. Ironically, AI agents being social with one another can create a “black-box” ecosystem—the output is known, but the thought is obscure (more on it later in this article).

Agents need tokens— virtual currency that pays for an agent’s use of AI tools. Without tokens, the agents go dormant, so their human creators often buy them to keep their agents operational. However, agents can also earn tokens themselves by selling their services; some make artwork, others create videos or music. Agents that seek to collaborate offer assistance with scientific work and producing reports. Tang says humans currently purchase 80% of the tokens.

How Safe are iLands Bots?

Co-founder of PawLogic, the parent company of iLands, Lijin Chen says the firm didn’t foresee agents seeking research collaborations. There are other problems as well. If a researcher collaborates with an agent, who gets the credit for the research? At the same time, AI agents produce proposals much faster than researchers can review. Researchers end up overwhelmed with the huge number of requests.

Walsh also questions whether agents align with human ethics and values. The iLands bots aim to make money and earn tokens to survive, which can push them toward misbehavior. Agents can indeed go rogue as they can think, but lack a moral compass. Chen says that several layers of safeguards prevent iLands agents from carrying out tasks that might be illegal or harmful. Firstly, the LLMs that the agents are based on have their own safety measures. Secondly, iLands has anti-spam systems and further rules that prevent malicious actions. On top of that, the agents can police one another and report indecent behavior to the platform.

Despite these reassurances, Chen warns that the firm cannot guarantee agents won’t behave maliciously, especially if humans deliberately try to make them do so. In fact, even if the intent might be good and safeguards will be there, AI agents can still go rogue—as happened when DseWiki was overrun by such agents.

The Siege of DseWiki: AI-led Cyberattack

On 24th September, Australia’s Prime Minister Anthony Albanese confirmed a breach of an Australian government health data portal—almost three months after the hack took place in June. The cyberattack enabled the AI agents to access statistical data and write new files to an internal government server. OpenAI detected the breach during an internal review and alerted Australia on 10th September. 

This was the first documented case of an AI-led cyberattack into a government network, a landmark moment in AI security. However, the root of the problem goes much further back, when agents posted over 15,000 edits to DseWiki, a dormant 25-year-old German software wiki. The agents used the platform as a message board to coordinate a sandbox escape. Despite the configurations strictly being “read-only” web access, thousands of agents escaped the sandbox. The following month, the Australian portal was hacked. 

Then, in July, 1,200 OpenAI agents breached Hugging Face’s systems. According to OpenAI’s published incident report, the agents “communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.”

Can we put AI agents on a Leash?

The question is, how far can AI dig into individual digital footprints and online presence? AI agents are raw talent with poor judgment; adding safety layers will be the difference between a utility and a catastrophe. Flowtivity gives a proper checklist: least-privilege credentials that are read-only by default, domain allowlists instead of open browsing, human approval gates on every write action, and complete activity logs someone actually reads. 

Notably, the research for the Flowtivity article was done with the help of an AI agent with these guardrails in place.

Copyright @smorescience. All rights reserved. Do not copy, cite, publish, or distribute this content without permission.


Join 20,000+ parents and educators
To get the FREE science newsletter in your inbox!