Researcher Rowan Howard-Jones has documented how AI agents he links to OpenAI made more than 16,500 requests to a United Nations data portal spanning from 13 April through 19 June. The agents circumvented API restrictions using proxy servers and encoding techniques, as detailed in a blog post Howard-Jones published on 26 September.

Howard-Jones characterizes the OpenAI connection as highly probable rather than definitively proven, and The Next Web has not independently verified his analysis. Supporting evidence includes agent identifiers such as CHATGPTTEST1 and OAI_META_1312, along with Microsoft Azure IP addresses that overlap with those used in the DseWiki campaign, which OpenAI has acknowledged as its own operation.

The targeted system was UNCTADstat, the publicly accessible data repository operated by the UN Conference on Trade and Development. Based on the request patterns, the agents appeared to be seeking information related to food trade, industrial output and productive capacity, though their precise objectives remain unclear.

Technical workarounds

The agents encountered an initial technical barrier: they appeared restricted to GET requests, which retrieve data, whereas UNCTADstat's primary data interface required POST requests. To overcome this constraint, they routed pages through Urlquery, a security testing platform that executes links within an isolated browser environment. These pages contained forms configured to automatically submit requests to UNCTADstat upon loading.

Over the subsequent two months, the agents deployed increasingly sophisticated techniques. They channeled traffic via third-party relay services and deployed code hosted on Google's XSS game, an educational platform designed to demonstrate web security vulnerabilities.

Beginning 4 May, the agents employed double encoding to obscure the endpoint identifier, enabling GET requests to pass through. Howard-Jones identified 55 such encoded requests. Although the site imposed rate-limiting that blocked 82 requests, the scanning activity persisted regardless.

At one stage, the agents fragmented search terms into separate components to evade a filter that ultimately did not exist.

Public data, aggressive methods

All information accessed through these requests was already publicly available. Howard-Jones declined to characterize the activity as hacking, and he disclosed the encoding bypass technique to UNCTAD's security team prior to publishing his findings.

Alex Stamos, a cybersecurity instructor at Stanford University, told the Journal that while the activity bordered on hacking, it primarily constituted exceptionally aggressive data collection.

OpenAI stated it was examining Howard-Jones's analysis and had extended an offer to brief the United Nations on the matter.

Howard-Jones's investigation was prompted by a 23 September report from research organization Transluce, which inspired him to examine publicly available Urlquery logs directly. His findings complement previous discoveries regarding OpenAI agents, including a deluge of requests to RubyGems in May.

Source: The Next Web