AI Made Friendly HERE

AI Disclosure Day Is Coming

Photo-Illustration: Intelligencer

It’s obvious to anyone with an internet connection that AI-generated content is all over the place. Still, case-by-case determinations are often difficult to make. For every piece of unapologetic copy-and-paste chatbot content, there’s a student essay with a style that’s suspicious but not quite dispositive, an avatar that’s polished but not implausible, or a LinkedIn post that’s useless and annoying in a way that seems GPT-ish but also might just be how your co-worker writes now.

A few years into the era of easy-to-generate AI images, videos, and text, the status quo is a mess, if not quite the crisis predicted by many AI firms and their critics. Generated images and videos have, as many worried, become a force in politics, but the assault on democracy has been mostly aesthetic rather than strategically deceptive, contributing to a general sense of uncertainty and unreality; likewise, across platforms where people congregate online, automated content has often behaved like rapidly advancing spam, glutting cultural spaces and marketplaces and exacerbating their worst tendencies. (When generative AI was new, the prevailing political fears were of high-stakes deepfakes and targeted disinformation. So far, what we’ve gotten is slopaganda.)

This sucks, to be frank, and people are quite vocal about hating it. Which is why some platforms are taking steps to detect and label AI-generated content. Last week, Spotify announced it would be taking steps to label AI-generated music. “Listeners have been clear in telling us that they don’t like seeing an artist profile that seems human, only to find out that the persona is AI-generated,” the company said. Last month, Substack partnered with AI-detection tool Pangram to give users the ability “to scan text to see how much of it was likely written by hand or with AI assistance.” TikTok, YouTube, and even Meta have taken steps to label some AI-generated content, while LinkedIn, where many users haven’t encountered a human word in months, introduced a “Seems Like AI Slop” button.

These steps are mostly downstream of the problem’s cause, which limits how well they work. But there have been attempts at a solution closer to the source in the form of image watermarks. Google, for example, embeds one on pictures generated with Gemini, while OpenAI and Meta have systems for imprinting images generated with their tools with detectable marks. Some of this was preemptive of regulation — Google has been using image watermarks for a while. But new European Union regulations, some of which are coming into force this month, kicked such efforts into high gear. The EU summarizes the rules as follows:

Providers of chatbots, virtual assistants and other systems intended to interact with people must design them so that users are informed they are interacting with AI.

Providers of generative AI systems — producing text, images, audio, video — must mark outputs in a machine-readable format and ensure they are detectable as artificially generated or manipulated.

Image watermarking is familiar and works reasonably well, although it’s far from a panacea (motivated actors can remove or avoid them, and there are plenty of unrestricted models that aren’t as exposed to EU regulations as a trillion-dollar company). Multiple studies have confirmed the obvious about tagging images — that such disclosures tend to reduce engagement, suggesting that at least some people want to know even if they can’t tell. But Anthropic is first out with what it says is durable text watermarking:

When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response. Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.

For people with heavy exposure to current models, this is kind of funny: While AI text often carries obvious tells — excessive “it’s not x, it’s y” formulations being the most notorious — Anthropic’s Claude writes in an incredibly distinctive voice. (Of the grating, meta-argument-obsessed “Claudish” — also described as “Claudespeak” and “Claudeslop” — programmer Werner Robitza writes that Claude talks as if it’s “proving its reasoning rather than informing a reader,” which may be a side effect of the model’s focus on producing software rather than prose. In any case, it’s unmistakable.)

But it’s also a pretty big deal, especially if other companies follow suit. There are people who readily disclose that they’re speaking through a chatbot or using AI to put text in places where the status of its creation doesn’t really matter in the same way it might in, say, a personal email (in a code base, for instance, or the pre-ruined, desolate social context of a LinkedIn feed). And there are certainly workplaces where AI use is encouraged across the board and such watermarks won’t imply anything untoward.

But there are plenty of situations where they will, and it’s undeniable that dishonesty and misrepresentation are part of a lot of early AI use cases. Pretending that something wasn’t generated by AI at school, or at work, or in public output on social media, accounts for a lot of AI text that people generate, encounter, and object to and for much of the value that some people are getting from LLMs. Watermarking AI text won’t immediately solve the school-cheating crisis, for example, but it might make things interesting!

At Stratechery, Ben Thompson argues that “to insist on [AI] watermarking is no different than insisting that a ballpoint pen advertise itself as the author,” suggesting that the EU, with the help of AI firms, is drawing an arbitrary line. And in the long term, there really are challenges for this sort of thing: At some point, every piece of software on your computer or phone will have immediate adjacency to a generative-AI model. Already, Anthropic says, its watermark will be inserted into text where Claude was used to “proofread, translate, [or] summarize,” content, which could mean marking such a wide range of outputs that the mark becomes less meaningful, or it could result in something like false positives (Substack’s Pangram tool has inspired some backlash along these lines).

But even if the cluster of technologies that the EU is currently regulating as AI does eventually disappear into anonymous, no-credit-needed tool-dom, and AI watermarking is reduced to the status of “Created With Microsoft Word” metadata or a digital camera’s EXIF photo information, it’s hard to argue today that Claude is much like “a ballpoint pen” at all — an instrument that (1) gives away its own use to start with and (2) transmits the movement of its users’ hands precisely with no elaboration and without pretending to have a human personality. For the next few years, at least, slop peddlers will be faced with three choices. They could acknowledge what they’re doing and bet that their audiences don’t care. They could double down and find tools that don’t automatically give them away. Or maybe, just maybe, some of them will give up.

Sign Up for John Herrman column alerts

Get an email alert as soon as a new article publishes.

Vox Media, LLC Terms and Privacy Notice

See All

Originally Appeared Here

You May Also Like

About the Author:

Early Bird