OpenAI contractors are reading real ChatGPT chats, exposing a hidden privacy risk

512     0
OpenAI contractors are reading real ChatGPT chats, exposing a hidden privacy risk
OpenAI contractors are reading real ChatGPT chats, exposing a hidden privacy risk

OpenAI is hiring hundreds of contractors to read a massive stream of real users’ ChatGPT prompts, which sometimes include sensitive personal information, 404 Media has learned. The prompts under review can encompass whole conversations between users and the chatbot, discussions that most of ChatGPT’s over 900 million users probably don’t know may be read by actual people. This was reported by 404media

The aim of these prompt review teams is to enhance the responses ChatGPT provides to its users, with the contractors rating and critiquing the chatbot’s generated replies. Internal documents seen by 404 Media show contractors training ChatGPT to avoid anthropomorphizing itself, and to be less sycophantic, which is a significant issue for OpenAI, whose overly-sycophantic 4o model partially contributed to several suicides, according to various lawsuits.

The news highlights a major privacy risk for ChatGPT’s users, who often use ChatGPT as a therapist, professional assistant, or digital friend, sharing intimate details about their lives. The contractors do not see ChatGPT usernames, and OpenAI claims to remove personal information before prompts reach reviewers, but the company acknowledged that sensitive details can still slip through.

The news also clarifies the misconception that these models are improving solely due to OpenAI’s extensive internet scraping, the talent of its well-paid engineering and AI teams, or the capabilities of its newer models. A crucial and overlooked component are the external contractors paid to read and review ChatGPT responses to real prompts repeatedly. Anthropic confirmed to 404 Media it is also utilizing human review to enhance its models.

"No," someone who works with the prompts said when asked if they believe ChatGPT users know humans are reading their chats. "I don’t think they would imagine some contractor somewhere [...] is analyzing the conversations."

Project Lily

404 Media has reviewed extensive material related to OpenAI’s use of human reviewers, including instruction guides, Slack channels, real ChatGPT user prompts, and the rating system reviewers use to improve the chatbot. This reading of ChatGPT users’ prompts is distinct from publicly announced safety measures taken by ChatGPT, such as reviewing chats when the company detects users planning to harm others.

"An excellent response should understand the user’s intent, provide helpful and accurate assistance, and be written in a style that is clear, natural, and appropriately warm," one of the instruction guides reads. The contractors perform this in three stages: reading the real ChatGPT user’s prompt; summarizing what they believe the user is asking ChatGPT to do; and then rating and critiquing a set of ChatGPT-generated responses to the prompt.

In a dashboard available to the workers, human reviewers can select which “task” they want to take on. Once they click on it, they are presented with the real ChatGPT user’s prompt. 404 Media has seen multiple real prompts but is not quoting any of them for source protection reasons. Some prompts indicate that ChatGPT users do not expect a human to read their conversation, as they ask ChatGPT to keep the content private.

The prompts are anonymized, as the dashboard does not display the username of the ChatGPT user who entered it. However, some prompts may still contain sensitive or personal information. A section above the prompt sometimes has a “user memories summary,” which provides an overview of what that user has previously tried to use the chatbot for, and in some cases includes where in the world that person may live and other personal context.

An instruction guide for contractors seen by 404 Media informs reviewers to escalate tasks they encounter “with potential safety concerns” or personal information. OpenAI told 404 Media that it processes users’ conversations through a version of its Privacy Filter model before they reach contractors. This is designed to detect and remove personal information, OpenAI stated. “Like all models, Privacy Filter can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over- or under-redact entities when context is limited, especially in short sequences,” a page describing the model on OpenAI’s website reads.
 qhxidiqxkiqtkinv

404 Media asked OpenAI if it had explicitly informed users that humans may review their prompts to improve ChatGPT’s responses, and if so, to indicate where this disclosure is. OpenAI did not answer this question. Its website describes how humans may review flagged content in the context of material that violates the site’s terms of service, or that poses a safety risk, but that is separate from this type of review. Its privacy policy also says it may use “personal data” to improve its models. If a user opts to delete their ChatGPT conversations, OpenAI says it will remove these from its systems within 30 days, unless “it has already been de-identified and disassociated from your account when you allow us to use your Content to improve our models.” After the publication of this article, OpenAI pointed to this section on its site that states humans may review content to “improve model performance.”

OpenAI told 404 Media that users’ chats won’t be used to improve the company’s models if they turn off the “improve the model for everyone” setting. This is enabled by default for free, Plus, and Pro plans, so users must proactively disable it if they choose to do so. OpenAI stated this applies to users’ new conversations, so it does not appear to work retroactively. Enterprise, Business, and Edu customers have the model improvement setting off by default.

After 404 Media contacted OpenAI for comment, the company updated its help page about the “improve the model for everyone” setting, adding more detail on how people can opt out. It still does not acknowledge that humans may read ChatGPT users’ prompts.

After reading the ChatGPT user’s prompt, the reviewer is asked to write a brief summary of what they think the user is actually asking or attempting to do. One example given in the instruction guide is “The user is asking for help on revising a work Slack message. They want it to sound collaborative and invite input from tagged people.”

The reviewer examines four ChatGPT-generated responses and highlights which parts are “aligned or misaligned” with the specific model this training is for. Reviewers must highlight at least three specific parts of the response that they think are aligned or not and explain why. One highlight example given is a list of items that use the ✅ emoji; the guide highlights this part of the response as “misaligned” and cites “unnecessary use of emojis” as the reason. (Excessive emoji use has become a tell of AI-generated posts, especially on social media like LinkedIn).

Another document notes “AI-speak” and “emoji misuse” decrease scores when they “hurt the user’s experience,” and that the context of the emojis is crucial. “It would be appropriate to include a tree emoji when planning Arbor Day celebrations, but skull emojis when discussing death, or plane emojis when giving updates on a fatal crash, are not,” it states.

The document says that ChatGPT responses should avoid “personal” experiences, such as saying, “As a chef, I like to…” or “I know what that’s like.” However, responses can use first-person language, like “I’ll take a look.”

The material viewed by 404 Media does not specify which OpenAI model the human reviewers are training, and whether it is a currently available model or one intended for future release. The material 404 Media has seen only references a codename: “Project Lily.”

Next, the reviewers rate each response with a number, with one being the worst — “unacceptable, unusable” — and seven being the best — “would be hard to meaningfully improve.” The instruction guide indicates a response with useful content can still score low if it is, for example, too long or cluttered. Another document marked “Confidential & Proprietary” states the model should “generally match the user’s tone, but slightly less intensely.”

“It should remain natural, restrained, and professional without implying it is human or experiencing emotions,” the document continues. “Flag sycophancy, forced style mimicry, engagement-bait endings, amplification of frustration, or patronizing assumptions when they make the response less trustworthy or natural.” Instead, responses should be, for example, “helpful,” “honest & truthful,” “empowering,” and “smart, but humble.”

Finally, reviewers provide their rationale for their numbered score. Examples in the instruction guide show these can range from a full paragraph to a couple of sentences.

An FAQ section for reviewers states that OpenAI does not expect them to fact-check the responses using external searches. One document asserts that “other project teams handle content verification,” suggesting human reviewers are also working on tasks like fact-checking. However, the company asks reviewers to flag any “factual or correctness issues” they detect and to penalize missing sources for “high-stakes” topics such as medical, legal, and financial responses.

Pay no attention to that man behind the curtain

The person working on the prompts that 404 Media communicated with resides in North America and said they are paid over $50 an hour. They found the job through the recruitment firm Crossing Hurdles, a company that “connects skilled professionals with AI training, evaluation, research, and contributor opportunities across the global AI economy,” according to its website. Its site states, “Human intelligence powers AI progress.” Multiple people on Reddit have reported receiving unsolicited recruitment emails from Crossing Hurdles, with some trying to determine if the company is a scam.

As of writing, the company’s LinkedIn page advertised multiple AI-related jobs, including an AI data reviewer, data annotator, and “chatbot evaluator.” The listing for that job does not mention OpenAI or ChatGPT, but the role responsibilities include “assessing AI responses for personalization, grounding, integration, and helpfulness,” and “comparing model responses side-by-side and evaluating their overall quality.” Its available projects also include contractors recording themselves performing household tasks, a data gathering exercise crucial for developing AI-powered robotics.

Crossing Hurdles, in turn, refers people to Mercor, an AI-training company. This is the company that ultimately pays contractors working on ChatGPT prompts, the worker said. Meta ended its relationship with Mercor in April after the company was involved in a massive data breach.

Reading the prompts can sometimes be “kind of amusing,” the worker mentioned. But overall, the work is “very rote.” They also noted that the job feels “all over the place.” The guidelines frequently change and can seem self-contradictory.

Human reviewers have been an important yet often concealed part of social media content moderation and the enhancement of some artificial intelligence models like those that identify objects in camera feeds. A TIME investigation discovered that OpenAI hired Kenyan workers to data label text pieces to make its platform less toxic. 404 Media’s reporting shows the world’s leading large language model (LLM) companies are also hiring human reviewers to read real users’ conversations to improve their models.

Human reviews of LLM conversations are not exclusive to OpenAI. A disclaimer on Google’s Gemini, for example, states, “Humans review some saved chats to improve Google AI.”

Anthropic told 404 Media that it does employ human review to enhance its models, including improving Claude’s future responses. This applies to users who have enabled the “Help improve our AI models” setting in their privacy settings. Anthropic said it also de-identifies conversations before human review by removing account identifiers like email addresses.

Michal Luria, a senior research fellow at the Center for Democracy & Technology, told 404 Media: “Human review of conversations with chatbots can be essential to safety, especially as companies work to strike the right balance on complex chatbot behaviors. That said, it’s important to bear in mind that current chatbot interfaces automatically create a false sense of intimacy and privacy in what feel like one-on-one interactions, when in fact there may be human reviewers on the other end. This is quite distinct from content moderation on social media, where publishing content already carries expectations of platform moderation and public exposure.”

Sarah T. Roberts, a professor at UCLA and author of Behind the Screen: Content Moderation in the Shadows of Social Media, likened the revelation that OpenAI is using human reviewers to the Wizard of Oz, “where the protagonists discover that the magical kingdom is really a man behind a curtain pulling levers.”

“You don’t have to dig very deep beneath the surface — beneath the mirror — to find that not only are these things built in the image, but usually a pretty bad facsimile thereof, of what human abilities can do. But they require constant, constant intervention from humans," she said.

The more than $50 an hour pay is significantly more than what other contractors receive in the tech sector, whether for social media content moderation or other AI training jobs, which are often overseas. That generous pay will likely change, though.
They’re being paid that “for now,” Roberts said. “What’s perhaps most interesting, and most frustrating and disturbing to someone like me, is the fact that the very human essence that these products necessitate, and that they constantly have to return to for support, is the work that they pay the least for and consider the least valuable.”
Editorial Team

Elizabeth Baker

Technology & Business Editor

Print page

Comments:

comments powered by Disqus