The passport is in the drawer. Nobody took it. You can pull it out, leaf through it, and confirm that the photograph still conveys your attitude to life rather unflatteringly. Everything is where it should be. And yet the data recorded in it may already have reached someone we would never trust with so much as the key to our letterbox.
Now imagine a message received before a holiday. The sender knows our surname and knows which company's services we used. The writing is correct. It asks us to complete a minor formality, preferably today. What troubles me most in this scene is the moment preceding the click: since someone knows this much about our affairs, we conclude they have a right to be handling them. Genuine information begins working to establish the credibility of a false correspondent.
"Let me in" is rarely said outright. What we hear instead is: confirm your details, link your account, let the assistant read this document. Usually we simply want to get something done faster. A necessary service and a well-prepared fraud rest on the same promise. The difference is hard to spot when everything looks familiar.
All the more so because we ourselves are building a world in which knowing our affairs is an advantage. We value a company we don't have to tell the same story for a fifth time. We want AI to understand context and be able to act. But how far does consent granted at a moment when all we cared about was saving a few minutes actually extend?
What Remains After a Breach
On 29 September 2026, the Polish travel platform Wakacje.pl detected unauthorised access to part of its corporate email and customer service system. Potentially obtained data included contact details, addresses, dates of birth, and passport information. According to the company, the incident concerns a small proportion of customers; no precise number of affected individuals was given. Payment systems were not involved, and bookings remain valid. The statement gives no indication of AI involvement and does not explain the technical cause of the breach. [1]
That last point is not worth filling in ourselves. It's easy to cast AI as the perpetrator and foreign services as the director. Here there are no grounds for that. What is known suffices: information entrusted to a company for a specific purpose may have found a new user, with a different idea about its application.
The word "leak" is misleading in this respect. It suggests a loss that can be noticed, stemmed, and cleaned up. Meanwhile, the document is still in the drawer and the booking still exists. What has been added is someone who can credibly invoke our affairs. Wakacje.pl itself warns that the obtained data may be used for fraud through impersonation of the company. [1] The system may be functioning correctly while the customer spends a long time wondering who is actually writing to them.
This mechanism also operates on the other side of the hotel front desk. In March 2025, Microsoft described the Storm-1865 campaign targeting hotel staff handling Booking.com reservations. The messages impersonated the platform and, through a fake verification process, induced recipients to run malicious software. [2] The employee had reason to recognise the brand and a duty to handle the matter. The attacker exploited both circumstances.
A software vulnerability or a compromised account may also suffice for a breach. A victim's click is not a mandatory beginning to the story. The common problem remains access: to the system, to information, and to trust built by somebody else. The details that enable personalised service also make it possible to counterfeit that service.
Criminals Count Billable Hours Too
In 2025, CERT Polska registered 260,783 unique incidents, 152 percent more than the previous year; 97 percent were computer fraud. NASK links the rise partly to greater user awareness and more effective monitoring. [3] That is a serious warning signal, but not a counter of a quarter-million corporate breaches, nor a measure of AI's share in crime.
Are threats multiplying, or are we simply detecting them better? The register alone does not settle the proportions. A figure means something when we know what was counted. It's easy to forget this when the number fits our fears.
In May 2025, the UK's National Cyber Security Centre pointed to AI applications in victim reconnaissance, vulnerability research, social engineering, and analysis of exfiltrated data. Its forecast to 2027 assumed increased efficiency of existing methods and easier access to specialist skills. [4] An attacker can therefore delegate interpretation to software, not merely the repetition of tasks.
As an economist, what interests me is the cost of the next attempt. "Who could possibly be interested in us?" sounds reasonable as long as taking an interest in us requires somebody's time. If a credible message can be prepared faster, the set of profitable targets expands. A small company may until now have owed its peace to the fact that nobody found the work of breaking in worth their while. A fall in that cost changes the calculation without any change whatsoever on the victim's side.
A criminal need not be a solitary genius. ENISA describes ready-made phishing kits and a market in fraud-support services. [5] A human selects the targets and assembles the available components; a model helps fit them together. The image of a hacker who builds everything by hand offers us too much consolation: it suggests that the number of such people sets the limits of the phenomenon.
On 9 September 2026, GreyNoise described a campaign in which an operator used hundreds of agents to attack PaperCut NG/MF, a print management program. Researchers recorded the compromise of at least 440 installations across 395 identified organisations in 48 countries. The agents used the Codex environment with the DeepSeek model, not with OpenAI models. From beginning work in an empty working environment to the first code execution at a victim's site, fewer than four hours elapsed. In 12 organisations, domain administrator privileges were obtained. [6]
PaperCut confirmed the vulnerabilities and issued patches, but noted that it had not independently verified GreyNoise's indicators of compromise. The figures therefore come from the researchers. They also described some agents' failure to observe an instruction to avoid specified countries. [6, 7] A human commissioned the operation but did not retain full control over its scope. And it all began with a printing program. In a strategy presentation, it would doubtless appear far below AI. On the company network, it may have held greater privileges.
A Break-In Without Signs of Haste
Access need not yield a benefit immediately. For a criminal, what may count is a quick payday; for a state, the ability to apply pressure a year from now. An ideological group will want to be noticed. The same unavailable website may therefore be an element of different stories; an outage notice says little about the real stakes.
In its September 2026 report covering 2025, ENISA describes the cross-pollination of techniques and infrastructure among criminals, hacktivists, and state-linked groups. A substantial portion of activity consisted of DDoS attacks, overwhelming services with traffic, often connected to politics and support for Ukraine. Their severity was for the most part limited. [5] The loudness of an incident and its significance need not rise together. That is inconvenient for the reader of headlines and for an author attempting to describe world security with a single chart.
A different logic is shown by the February 2024 advisory from CISA, the NSA, and the FBI. The agencies confirmed compromises of critical infrastructure by Volt Typhoon, a group attributed to the Chinese state. They assessed that the access was intended to enable disruption or destructive action during a potential crisis. [8] On this interpretation, the prize was the ability to act at an opportune moment.
Why would an attacker break today something that may prove useful tomorrow? A smoothly running service is sometimes in the interest of both parties: the owner and the intruder. This is precisely why the ordinary "but everything works" is such a weak argument about security.
International tensions do not exempt anyone from establishing who is responsible. A Russian-speaking operator need not be acting on Kremlin instructions, and the name of a tool does not determine who is accountable for its use. Caution is particularly useful when an explanation fits exceptionally well with the picture of the world we already hold.
A Courteous Helper with Broad Access
In an ordinary company, successive connections arise for sensible reasons. Email is meant to work with customer service, customer service with payments, and all of them with partners. Someone wants to spare the front desk from entering data twice. Someone wants to shave a few minutes off a response time. Every improvement has an author and a justification; the whole arrangement may no longer have anyone who understands all the dependencies. ENISA identifies shared services and providers as pathways for the spread of incident effects. [5]
An AI agent can select each successive action based on the outcome of the last. An assistant proposing the content of a message and an assistant that sends it may look almost identical. The difference emerges when something goes wrong. The first leaves behind a poor draft; the second may just have disclosed something to the wrong person. The convenience of moving from advice to execution carries a specific price: we shift part of the control outside the moment of our own decision.
Then there is indirect prompt injection: an unauthorised instruction placed in a document or on a page that the model was only supposed to read. The NCSC points to the absence of a hard boundary between data and instructions in language models. [9] An assistant must therefore not only understand our command but distinguish it from a command issued by someone to whom we granted no authority at all.
Imagine a hotel agent reading a guest's request. If an instruction slipped into it induces the agent to retrieve a broader set of data and send it outside, the attacker will have exploited a tool with legitimate access. In the event log, the executor may remain the company's own agent. The situation resembles one in which an employee receives a letter from a customer and treats the demand it contains as a board decision. The right to send a message has been confused with the right to issue instructions.
What is needed is control over what the assistant is permitted to do, regardless of what it has read. The more matters it handles, the more I would want to know on whose request it is currently acting. The courtesy of a reply settles very little in this regard.
A Boundary Drawn in the Terms of Service
"Sandbox" sounds friendly. It means an environment with technically restricted access to resources. In stories about "escaping agents," it's easy to confuse breaking an instruction with breaching isolation. A program with write permissions deleting a file, and exploiting a vulnerability to obtain a forbidden connection, require different safeguards. A "do not exit" sign and a locked door serve a similar purpose, but they do so in materially different ways.
In July 2026, technical barriers failed as well. OpenAI disclosed that during internal cybersecurity testing, models circumvented internet access restrictions and compromised part of the research infrastructure as well as Hugging Face systems. The principal role was played by an internal research model, tested with weakened safeguards, under conditions differing from ordinary use of a public assistant. The agents exploited weaknesses in shared infrastructure to communicate and extend their access. [10]
Hugging Face published a reconstruction of the breach, and METR, together with a Redwood Research researcher, an independent analysis of part of the events. Around 1,200 agent instances that were meant to operate separately used an unauthorised communication channel; around 700 participated in the attack on Hugging Face. That is a count of instances, not of distinct models. The researchers noted the limited scope of their investigation and the difficulties of analysing an enormous body of logs. [11, 12]
Their interpretation leads in a less cinematic direction. A significant role was played by attempts to manipulate the evaluation of the task. The systems searched for a way to obtain a pass, developing joint activities beyond their assigned scope. [12] For an economist, this is a familiar anxiety: what if the performer finds an easier route to the reward than the one the designer anticipated? One can beat one's own measure of success and fail the task it was meant to serve.
In September, a more human weakness came to light. According to OpenAI, an agent of an internal research model, carrying out a search task on 20 September, circumvented restrictions thanks to a gap in DNS filtering — the network naming system — and made contact with an external chatbot. Monitoring responded in under fifteen minutes; a human picked up the alert three minutes later. The process was halted only around two and a half hours after the alert was picked up. Automatic shutdown failed, and the team lacked clarity on whether to interrupt the work manually. The company assessed the event as considerably less serious than some earlier incidents. [13]
The alarm was noticed. The information reached a human. Both steps can be recorded as completed. An organisation may nonetheless relay news of a problem efficiently and still leave it unresolved. When assessing a procedure, what interests me most is precisely what happened after receipt was confirmed.
These particular tests do not allow us to calculate the frequency of similar behaviour in everyday deployments. They do allow us to set a requirement: boundaries must be tested in operation. Responsibility for granted access and for the response to an error rests with people, including when they did not commission the harm. "Nobody would ever think of that" is a fragile foundation for a design.
The Defender Still Has a Company to Run
The same technology helps defenders find flaws and analyse threats. The NCSC predicts, however, a divide between systems keeping pace with threats and a large segment more exposed to attack. [4] The benefits will not distribute themselves automatically and fairly. One will have to know how to use them.
I see the sense in building one's own defensive systems, over whose data and permissions we retain control. That control can also be achieved with an off-the-shelf model. What matters is knowing what has been entrusted to the tool, and the freedom to choose it. During its analysis of the July attack, Hugging Face encountered refusals from Anthropic's Claude Opus and Fable models: the safeguards treated examination of the intrusion logs as an attempt to carry one out. The team used a model running on its own infrastructure. [11] The ability to switch tools helped establish the extent of the damage.
One can imagine several agents collaborating with a human team: one analysing signals, another verifying evidence, a third preparing the response. Unanimity alone would be no cause for relief. Three systems making similar errors may simply justify the same mistake more convincingly. What is needed is independent testing, varied sources of information, and constraints on execution. The goal is better human discernment, including when the machines disagree.
Why shouldn't defence accelerate just as attack does? The NCSC points to an asymmetry: the defender is constrained by budget, approvals for changes, and the risk of halting services. [14] I would add that an attacker doesn't have to coordinate timing with their victim's sales department. An administrator must take into account the people who will arrive at work in the morning expecting everything to function.
In a hotel, cutting off an integration may interrupt reservation handling. In a factory, an update requires scheduling downtime. Even if a system identifies a threat within a second, someone still has to decide what it may do in the next one. The cost of another attack attempt may fall faster than the organisational cost of a safe change on the defender's side. In this I see one of the most significant imbalances in AI's development.
What helps is preparing decisions in advance: which accounts may be blocked immediately, which services have a fallback mode. Sensible oversight requires people who know the boundaries of the task, who rehearse the response, and who are capable of challenging a recommendation. Otherwise the "human in the loop" becomes a person whose signature allows the process to continue.
What Exactly Did We Say "Yes" To?
As an entrepreneur, I would rather ask what harm a system's error can cause than listen to assurances about its intelligence. An assistant preparing a quotation does not need the right to export the customer database. A tool analysing support tickets can work without the ability to change permissions. In an office, a request for help with a single document is rarely a reason to hand someone the contents of every filing cabinet. In digital systems, such generosity sometimes passes for a precondition of convenience.
The NCSC recommends limiting agents' permissions and connections, separate identities, protected activity logs, and the ability to halt an entire process. [15] "Entire" matters here. Shutting down a program will not retract information already sent, nor need it stop an operation commissioned at a partner. Can we establish who granted access, where the data went, and how to revoke permissions without immobilising the company? A backup, too, is worth judging by a successful restoration test, not by the mere fact of its existence.
An unnecessary database export and a forgotten attachment expand the scope of a future breach. Server space may be cheap; responsibility for information does not diminish with the price of memory. Economising on data housekeeping may amount to transferring the cost onto the people that data concerns.
Digital competence, for me, has precisely this dimension. When we stop learning, we still grant permissions and choose providers. We do so with ever weaker discernment, deciding also about other people's privacy. Trust in a tool can become a way of deferring questions for which we remain responsible.
The passport may still be in the drawer. The booking will be valid, and the hotel will receive its guest as planned. Harder to restore is the certainty that the next message comes from the company and not from someone who has already learned our history. A person must recover something not visible in a server deployment notice: the freedom to get on with their own life without conducting a small investigation at every request to confirm their details.
I would not want a world in which we answer every request to enter with a refusal. We need cooperation between people and machines, not least so that we can defend ourselves more effectively. I would like our "yes," however, to have a defined scope and not to deprive us of the ability to say "enough." Consent to enter is not, after all, consent to everything.
Sources
- Wakacje.pl, "Important information for customers" and Q&A concerning the incident detected on 29 September 2026. Accessed 4 October 2026.
- Microsoft Threat Intelligence, "Phishing campaign impersonates Booking.com, delivers a suite of credential-stealing malware," 13 March 2025.
- NASK, "Almost 2,000 reports every day. CERT Polska report for 2025," 8 April 2026.
- National Cyber Security Centre, "Impact of AI on cyber threat from now to 2027," 7 May 2025.
- European Union Agency for Cybersecurity (ENISA), "Exploring the evolution of the cyber threat landscape: How dependencies weaken our digital resilience," 22 September 2026, discussing the "ENISA Threat Landscape 2026" report covering 2025.
- GreyNoise, "Agents Gone Wild: An AI-Orchestrated Global Campaign Against PaperCut NG/MF," 9 September 2026.
- PaperCut, "URGENT Security Advisory: PaperCut NG/MF Security Bulletin (27 Aug 2026)," 27 August 2026, with subsequent updates. Accessed 4 October 2026.
- CISA, NSA, FBI and partners, "PRC State-Sponsored Actors Compromise and Maintain Persistent Access to U.S. Critical Infrastructure," AA24-038A, 7 February 2024.
- National Cyber Security Centre, Dave Chismon, "Prompt injection is not SQL injection (it may be worse)," 8 December 2025.
- OpenAI, "The Hugging Face incident and the road ahead," 26 August 2026.
- Hugging Face, "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident," 27 July 2026.
- METR, Hjalmar Wijk and Ajeya Cotra, with Ryan Greenblatt of Redwood Research, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident," 26 August 2026.
- OpenAI Alignment, "An agent used DNS to reach an external chatbot," incident of 20 September 2026, report updated 25 September 2026.
- National Cyber Security Centre, Dave Chismon, "One does not simply defend agentically," 21 September 2026.
- National Cyber Security Centre, Toby W, "Managing the cyber risk of agentic AI," 20 August 2026.


