Rendered at 14:24:16 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Eddy_Viscosity2 3 minutes ago [-]
Is there a way to not only watermark that it was AI, but also which user generated it? Can it be used to fingerprint?
firefoxd 23 hours ago [-]
I feel like this is going to end up being like cookie laws. It sounds good, I don't know how any one benefits from it.
Currently I can recognize AI text because I read thousands of ai generated text. I know that 110% of yahoo finance news is generated. I don't want to read an AI generated personal blog, but if I do what's the problem really? Other than the companies distinguishing AI text for getting better training data, how do people benefit from watermarked text exactly?
tcdent 22 hours ago [-]
This feature is born out of the same government that created the cookie laws, so it makes sense to me that it would follow similar politically-valuable but practically-questionable set of beliefs.
Really, Anthropic (and the other labs that follow) are just trying to satisfy the requirements of the law so they can continue to serve the EU. Wether it's actually effective is something else entirely.
Lorkki 4 hours ago [-]
The cookie laws and GDPR are hampered by malicious compliance, sure, but also outright intentional non-compliance. What's needed is more efficient oversight and enforcement.
theoreticalmal 2 hours ago [-]
How would more efficient compliance help the constant annoying pop ups?
setopt 52 minutes ago [-]
According to the rules, accepting or refusing cookies should be equally difficult. Making them both difficult makes people not use your site, so the incentive should be to make both options easy if you follow the rules.
Malicious compliance is to make cookie acceptance much easier than refusal. Lack of oversight doesn’t punish this. Now, the incentive is to make life hard for people who decline cookies. That results in more annoying pop-ups.
ffsm8 32 minutes ago [-]
also, session cookies for eg login and error tracing etc dont actually need a banner. you only need to use it if youre participating in the internet stalking behaviour by integrating eg google analytics and similar integrations
theshrike79 7 hours ago [-]
There are stories of students literally copy-pasting LLM answers to school assignments with prompts and all.
This will most likely bring a sift end to at least the low hanging fruit.
mFixman 5 hours ago [-]
Removing AI watermarks is still easier than doing a school assignment.
The only way to prevent AI cheating is to make the assignments in class in pen and paper. To be honest, I don't understand why schools moved away from that in the first place.
graemep 5 hours ago [-]
> I don't understand why schools moved away from that in the first place.
Cost and convenience.
Being able to work on computers does help people with bad handwriting, which includes some disabilities.
anshorei 5 hours ago [-]
Yes, and back in my school days we all did our tests on paper, except the two kids who had a slip of paper that allowed them to take a digital test instead. Helping people with disabilities doesn't require having everyone make their tests on a computer.
And taking away writing from people with bad handwriting (if not related to disability) is just about the worst thing to do if you want them to ever get better at it.
theoreticalmal 2 hours ago [-]
Idk, I’ve had bad handwriting my entire childhood and I still have bad handwriting now as an adult. I’ve (somewhat jokingly) announced that I am “illegible”. It isn’t a detriment to my day to day activities at all, because I can type sufficiently well
mFixman 5 hours ago [-]
And now instead of teaching disabled students how to improve their handwriting to improve their lives regardless of their circumstances, we lowered the bar and made all students have bad handwriting.
graemep 5 hours ago [-]
I am talking about disabilities which affect things like motor control which means that they cannot improve their handwriting very much.
How does good handwriting improve people's live significantly? I know some people who prefer handwritten notes and I think there is research that says they can be more effective when learning, but most adults very rarely write more than a few lines by hand.
mFixman 4 hours ago [-]
You could say the same of PE, maths, literature, or almost anything learned in school. Few adults need them in their lives, but they provide good habits and virtue to students. People who don't read and write anything outside social media are a loss to themselves and to society.
Disabilities don't disappear after school, and children who struggle with motor control in school will continue struggling with it on adulthood. Half the point of school is finding out your weaknesses and learning to overcome them in a safe and controlled environment.
graemep 4 hours ago [-]
Studying maths and literature should develop your mind in many ways and gives you a chance to enrich your life. I do not see how handwriting will do that.
Pushing someone with a disability to improve their handwriting will not cure their disability. How does making people struggle more help them? its like suggesting kids in wheelchairs should be made to try to walk.
graemep 4 hours ago [-]
What I am not clear about is who will have access to the tools for identifying AI content? Will it be public? An API?
theshrike79 24 minutes ago [-]
AFAIK it's going to be a public API from Anthropic. Dunno about the costs.
Eddy_Viscosity2 3 hours ago [-]
Whoever pays for them most likely. Also likely that there will some special layer of watermarking capability only the sec agencies get to use (and pay more for).
conartist6 6 hours ago [-]
Should open up some jobs currently held by unscrupulous people as well ; )
Boltgolt 8 hours ago [-]
> I feel like this is going to end up being like cookie laws
Companies will implement it in the worst way possible? Nowhere in the cookie laws does it say you need to add a banner, you just can't spy on users without their consent.
LoganDark 7 hours ago [-]
The point of the banner has always been to malicious-compliance the users into being annoyed by the laws. Same as when one of the food delivery apps responded to a fair wages regulation by adding a special line item to every single order with help text that directed the user to appeal the regulation, it's just cookie banners usually don't say it explicitly
Joeri 7 hours ago [-]
The web’s ad markets bake in a loss of privacy by design, and so every website that wants to participate in those ad markets has to dupe as many users as possible into handing away their privacy. Couple that with declining web ad revenues because of the move to mobile apps, and you get a web filled with malicious compliance cookie banners.
Had google offered a privacy-preserving ad model (basing the ads off the content of the page instead of who the user is) things would be very different.
ifh-hn 7 hours ago [-]
I would income education institutions might benefit from being able to distinguish between AI and human generated text. The people who benefit, in the long run, are the students.
nonethewiser 15 hours ago [-]
>I feel like this is going to end up being like cookie laws. It sounds good, I don't know how any one benefits from it.
Does it sound good?
Either anthropic does not tell anyone the signal and only they can verify the text. Or they share the signal and everyone can verify it, which means everyone can bypass it and its just an inconvenience.
throwaway27448 9 hours ago [-]
Are they not offering an API to perform this check? Would be pretty simple to use any open model, even a weak local one, to basically tweak the text until it doesn't flag.
nradov 9 hours ago [-]
There are limits to what you can hide in text. If there is a signal then obviously independent cryptographers will be able to find it.
14 hours ago [-]
UltraSane 7 hours ago [-]
You can't think of the benefit of having LLM output watermarked? So many lazy students are going to get caught by this.
razakel 22 minutes ago [-]
Cheating was obviously going to happen when education became a business.
avaer 21 hours ago [-]
None of this is science or evidence based, it's political theater. The labs do this because the government has made it clear it will intervene if they don't say the line. I suspect most of the people at the labs think AI watermarking is a retarded approach.
The people that make the laws and the people complaining about AI don't know or care about the reality of the situation nearly as much as they care about being re-elected and feeling good about their social posture.
mschuster91 7 hours ago [-]
> None of this is science or evidence based, it's political theater. The labs do this because the government has made it clear it will intervene if they don't say the line. I suspect most of the people at the labs think AI watermarking is a retarded approach.
Well, the negative effects of AI are already visible, and I'm not talking about job losses here. I'm talking about AI hallucination, about mass generated spam, about AI-assisted scamming (apparently Indian scammers are shifting to use AI systems [1]) and about political manipulation.
Governments usually are slow to react, but eventually they do react, when a technology poses more issues than it is worth.
I for one value the transparency of being able to tell human-written content apart, especially when moving into a time where we’re suddenly questioning an AI agent's motivations. If that is ingrained so deeply into these systems they cannot rip it out without us noticing, there’s at least one more safeguard. I know how ludicrous this sounds, but you can’t deny AI development is speeding up to scary levels.
Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
nonethewiser 15 hours ago [-]
This doesn’t add up.
How do you think you will know the text was AI generated but the person trying to deceive you won’t be able to undo the watermark?
optionalsquid 14 hours ago [-]
Practically speaking, most people probably won't know how to undo the watermark. Especially when the suggested method is to run the text through a local model
cwillu 11 hours ago [-]
Most people are entirely capable of copy/pasting their text into the inevitable dozens of web-based tools of various levels of sketch to do it for them.
nonethewiser 14 hours ago [-]
Doesnt matter - it ruins integrity and the people actively trying to deceive you wont be prevented.
optionalsquid 14 hours ago [-]
What do you mean that it ruins integrity?
nonethewiser 13 hours ago [-]
I mean you cant be sure the text isnt AI.
nradov 9 hours ago [-]
I deny that there is anything scary about AI development. It's working fine for me so far.
josephg 15 hours ago [-]
Me too! And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students.
This article says any AI watermarking can be defeated trivially, by running the text through a tool which strips weird Unicode characters (fine) and then runs it through an LLM which replaces all words with similar words (whaaattt?). Running text through another - probably much lower quality llm which scrambles the words you use would make slop even sloppier. And it’s another whole step you have to know to do. A great many people who use LLMs to avoid doing work will not know about these extra steps, or not make use of these sort of tools.
nradov 9 hours ago [-]
Even if some of the commercial LLMs embed text watermarks, students will still be able to use open source LLMs with no watermarks. Universities need to stop goofing around and restructure every part of their curriculum around the assumption that students will use those tools. Trying to treat it as cheating is just pissing into the wind. Circumstances have fundamentally changed and there's no going back.
sib 14 hours ago [-]
>> And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students.
I would say that's not necessarily the case unless there are zero false positives. In fact, your university situation is exactly where a detection system that works 50-90% of the time would be a nightmare if a meaningful share of the 10-50% errors were false positives.
nonethewiser 13 hours ago [-]
And students could just verify their ai generated code doesnt have the mark.
Im very confused how this is even supposed to work at face value.
1) If the verification can be done by anyone, then anyone can bypass it.
2) If it can only be done by anthropic then the government or whomever has the special privilege (not everyone otherwise this is just #1) has to make a specific request
Is the point of the legislation to accurately classify text in general or to simply detect the true positives?
josephg 14 hours ago [-]
A watermarking system like this should have an incredibly low false positive rate. Orders of magnitude lower than the current crop of AI detector tools.
nonethewiser 13 hours ago [-]
Aka some kids are gonna get majorly fucked. Just not many.
josephg 13 hours ago [-]
The lazy kids, yeah. But, it was always the lazy kids who are cheating the most with LLMs.
EDIT: Sorry, I misread your comment above. Yeah, hopefully orders of magnitude fewer kids than the number who are getting falsely accused of cheating with LLMs now.
Increasing the accuracy of these systems - both in terms of false positives and false negatives - seems like a good thing.
rightbyte 8 hours ago [-]
Apple's CSAM hash filter removing peoples photos (did it notify the police too?) could be an indicator of how false positives on big scale work out.
josephg 7 hours ago [-]
How accurate is apple's CSAM scanner? 90%? 99%?
A proper watermarking system should be able to have an arbitrarily high accuracy - as many nines as you want. And it should be able to actually report the accuracy of its judgements.
If you're worried about kids being falsely accused of cheating using LLMs, you should be cheering on these developments.
If some kids are false positives (detected as using ai but didn’t)then how are they lazy?
vaylian 8 hours ago [-]
> Me too! And a detection system that works 50%-90% of the time is better in a lot of cases than no automated detection at all. For example, universities need to make cheating with chatgpt risky for students.
What about false positives? Imagine being a honest student and then the software declares your work to be AI-generated. How do you defend against that claim? The detection software is a black box and is likely running as a cloud service, so that you have no realistic options for reverse-engineering the false positive detection.
potsandpans 13 hours ago [-]
Now you just get to think, "did this person strip the stenography from the content?"
Nothing will save us. You can't automate trust.
giridharkannan 8 hours ago [-]
https://declaude.org/watermarking/ did a good job in explaining how SynthID works. As per their blog, it feels like it will be difficult to remove watermarking on bigger text and the checking for watermarking is also not complex
The article mentions you can always simply use a smaller, local, un-watermarked LLM to rephrase the original watermarked text. Which is true, sure.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
johnjwang 22 hours ago [-]
There exist methods to detect what kind of watermarking tool that someone is using, and most of the big tools have specific signatures that you can look for.
For Anthropic, it’s highly likely that the watermark is a SynthID type mark similar to the one that Sean is talking about (I actually ran the analysis here https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...). When we get confirmation of whether all models actually are watermarked, I think we’ll be even more confident.
Of course it’s always possible that Anthropic has come up with a proprietary scheme, but I think it’s definitely harder to implement.
I think the game will be a cat and mouse game similar to LinkedIn and other websites trying to block scrapers: each iteration makes it harder for someone to figure out the watermarking scheme, but likely not impossible
theshrike79 7 hours ago [-]
This always reminds me of the Trace Buster Buster Buster -scene from The Big Hit :D
>There exist methods to detect what kind of watermarking tool that someone is using, and most of the big tools have specific signatures that you can look for.
Serious question… what is the point if you cant verify the watermark? Then only the model provider will know.
Is that what this law is about? I thought it was so anyone would know (which of course means anyone can bypass).
gblargg 22 hours ago [-]
If you can submit the text to determine whether it's watermarked, you just have to progressively alter the content more and more until it passes.
rzmmm 2 hours ago [-]
This just shifts the legal responsibility to you.
nonethewiser 13 hours ago [-]
Yes this is what im trying to figure out. It means you cant be confident a negative is true. It sounds like this is the spirit of the law but it makes no sense. It would work better if only the model provider knows and the government can ask.
nonethewiser 13 hours ago [-]
>Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc.
Doesnt this and probably all techniques require the validator to know which portion of the text to validate?
If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm?
In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
recursivecaveat 21 hours ago [-]
I don't know if it's possible to achieve a watermark that is undetectable, has a low false positive rate, and survives a wholesale rephrasing. Can you make a statistical measure that reliably survives 95+% of the words being different and the sentences reordered? Of course the more of the content you replace, the lower the quality, but in many cases you probably care more about the meaning of the text than the exact choice of words.
Gigachad 15 hours ago [-]
Most AI sloperators can't even be bothered to remove the EM dashes or emoji spam from older models. The watermark doesn't have to be literally impossible to remove. If someone spent significant effort to rephrase it then let them have it. But the bulk of spam will be easier to detect.
nonethewiser 13 hours ago [-]
Well it depends on what you want to achieve.
Catch true positives? Sure
Reliably determine if text was generated by AI? Not even a little.
ProfessorLayton 20 hours ago [-]
>...and survives a wholesale rephrasing.
There's also literal language translation. Generate in language A, translate to language B (Either "manually" by being proficient in it, or with non-LLM translation).
A very, very significant portion of the world knows more than 1 language.
josephg 15 hours ago [-]
I think that would defeat the watermark. But low quality language translation tools have low quality output. If llm watermarks are defeated by making slop even sloppier, it’ll at least make it easier for humans to tell the difference.
And the more hoops you make cheaters jump through, the better.
ProfessorLayton 14 hours ago [-]
>But low quality language translation tools have low quality output.
Perhaps, but no tools needed for those who know more than one language, which is a lot of people.
josephg 14 hours ago [-]
Yeah, but you’re still manually translating a whole essay or design document or whatever between languages. People who get LLMs to do their work for them do it because they don’t want to spend time and effort writing. Making people do a translation process like this removes some of the benefit of using an llm in the first place.
Watermarking will never be a perfect tool. But there’s a lot of value in making low effort llm slop detectable. Even if high effort llm slop is still undetectable. Don’t make perfect the enemy of good.
cayleyh 23 hours ago [-]
It could be, but common, we all know that Anthropic's watermark is the using "load bearing", "genuine", and "seam" 1000x more in the same paragraph than any human in history.
queenkjuul 2 hours ago [-]
Worth surfacing that you're absolutely right
ljm 21 hours ago [-]
[dead]
Otterly99 2 hours ago [-]
I don't really see how this would work anyway.
Even if you find a way to 100% watermark any text, couldn't you just use a non-watermark model? I have a hard time believing every AI company on heart would comply.
Hell, even if every AI company on earth decide to somehow apply watermarking to their next model, they would also need to apply it to all the previous version that they commercialize. Given that anybody could make a copy of an open source model right now and would be safe forever, this seems quite the lost cause to me.
theshrike79 43 minutes ago [-]
This is just managing the low-effort low hanging fruit.
The students who, without a single neuron activating, will just copy paste questions to an AI web UI and paste the answers back.
Or the mid-tier corporate people who file 50 page "reports" that are 100% AI bullshit.
None of them will install a local model with no watermarking or bother looking up some Badonk AI service to do the same thing.
happytoexplain 23 hours ago [-]
Yeah, but it's better than nothing.
People underestimate the value of rules that only take malice and a little knowledge to break.
And they tend to exaggerate that underestimation if they... don't like the rule.
bonoboTP 22 hours ago [-]
I even see the Gemini diamond watermark on so many fake social media profile pictures. You could ask an LLM to find a github project that removes/inpaints those watermarks and be done with it. Or you could just use the API, which doesn't stamp the visible diamond on it (just the invisible synthID).
Most people are generally lazy. They upload text to LinkedIn full of "genuine", "honest" and "load-bearing".
josephg 15 hours ago [-]
> Most people are generally lazy.
And the people who get LLMs to make content for them are often the laziest of us all! Some people can’t even be bothered to remove their prompt from the output. Or they submit work which ends with “is there anything else you want to know about ___?”
Go ahead. Make tools to rip off the llm stenography. Many of the people who would proactively use tools like that are writing content by hand anyway.
nradov 9 hours ago [-]
How is this "better than nothing" considering that it's not even needed in the first place?
pavon 21 hours ago [-]
Yes! How many people have been caught committing fraud because they used a modern word processor and fonts to forge old documents? How many leaks have occurred because people failed to redact documents correctly, despite there being easy to use tools for this very purpose? How many people neglected to strip sensitive EXIF information from images they share (before websites started doing it for them)? How many people flat out post evidence of their crimes on social media?
Yes, this watermark will be easy to strip. It is still valuable for the vast majority of times where people just don't.
JoshTriplett 22 hours ago [-]
Exactly. If this works on pull requests, for instance, it'd be really useful for projects trying to do a first-pass filter to close slop spam.
hennell 21 hours ago [-]
Can it work on code itself? Obviously if you get AI to write the PR description that could be watermarked, but code isn't going to like random unicode, and the sythID approach feels like it would fall apart in the quite strict syntax of most code.
pavon 21 hours ago [-]
If a language supports unicode it will have a fairly permissive definition of whitespace, and it will be easy to generate permutations of the whitespace that meet the syntax requirements.
Of course I would want my code formatting tool to normalize that all to plain 0x20 spaces. But it would still be a helpful "brown M&M" test of did you even read CONTRIBUTING and run the code formatter before submitting this PR?
JoshTriplett 21 hours ago [-]
Code submitted to a project should follow the formatting standards of that project and its programming language, which doesn't leave much room for flexibility in whitespace. However, there are many ways to write comments, and many ways to write PR descriptions and commit messages, and often many choices of words to name identifiers, and other potential sources of bits of entropy with no functional impact. It doesn't take that many bits to encode a robust "AI was here" indicator.
krapp 19 hours ago [-]
It's worse than nothing because it gives people the false impression that there exists a system they can trust to tell them what information is authentic and what isn't. The incentive for governments and other interests to undermine that system to push false narratives and disinformation is just too great.
simon84 21 hours ago [-]
The goal of the AI act is not to determine if an "oh yeah!" comment was AI generated. The target is long papers that falsely claim human review and can have real significant consequences.
E.g. research paper, law makers, lawyers, state policies, notaries,...
These are much longer content and thus statistically they will disclose a better guess at AI generated content.
Asking another AI to paraphrase will not erase the mark (which they are unaware about) but rather cumulatively add their own mark and make it easier to detect.
The problem is not to use AI, but to endorse the responsibility of the content you (as a human) deliver and somehow make sure that fake-news, biased content or unverified output is detected as early as possible.
wodenokoto 10 hours ago [-]
I don’t understand how asking another AI to paraphrase causes both watermarks to remain
simon84 8 hours ago [-]
From what I understand of the not-very-detailed text watermark, it comes up to masking the generated output with specific bias in the weights of the token generation (or something like this). Anthropic was explaining that it would survive edits (not full rewrite). So asking another AI to paraphrase will surely add its own, but since the bias is not disclosed, it will depend on how the 2nd AI is considering the existing tokens to steer its own.
Fictive example: imagine the bias is to inject notion of colorfullness in the text regardless or its content:
"I eat an apple" becomes "I eat a red apple"
The 2nd AI is about injecting size component, so the text becomes " I eat a big red apple"
There you get both watermarks.
It is not that obvious obviously
avaer 21 hours ago [-]
You'd think that the humans being paid to review these things are actually reviewing them. Journals, laywers etc. are expensive. With the addition of AI, it should be easier than ever to review things on their merits.
Maybe some part of it is that the deluge of slop is uncovering how poorly/sloppily these social institutions were working in the first place.
maleldil 10 hours ago [-]
Peer reviews is largely done for free. In some cases, an author needs to commit to reviewing someone else's work for their work to be reviewed. However, this has led to an increase in LLM usage for review generation, even if conference/journal guidelines forbid it. I've seen nonsensical reviews from people who obviously haven't read the paper beyond an LLM-generated summary. Area chairs are supposed to catch this, but I imagine they're using LLMs too.
simon84 18 hours ago [-]
I think you are very correct here! Now imagine how much worse it can get with AI in the way.
It is as always (think cybersecurity) the cat and mouse game, what is a weapon is also a defense. You, the simple fact that you are on this very site, means that you are probably more educated to AI than most, so it is not necessarily you that will really benefit from any constraining framework.
There are many people believing in many weird theories and those are more prone to be convinced by a nice narrative. AI did not introduce that, it just made it easier and available to anyone with any intent.
watwut 9 hours ago [-]
> With the addition of AI, it should be easier than ever to review things on their merits.
It makes it harder, because amount of bullshit goes up and it is easier to generate plausibly sounding bullshitm
truckyday 21 hours ago [-]
[dead]
avaer 21 hours ago [-]
Anecdote: I've had at two friends publish "papers" just by slapping their names on something they had literally nothing to do with, out of pure nepotism.
This wouldn't be helped by an AI watermark. What would help is if the reviewer used AI to look up the authors and see the authorship claims are dubious. The papers are still up.
I do think the academic publishing field is corrupt, which is why I'm not convinced it should be on AI providers to help bail it out of doing its one job (verification and trust).
maleldil 10 hours ago [-]
Verifying authorship is an ambiguous task. How much involvement is necessary to qualify for authorship? If someone reads the paper and gives some small feedback, is that enough? What if they were present in one meeting and raised a question that turned out not to be interesting? What if they have no clue about the work but helped with data validation?
I'd argue all of these could justify authorship, even if they're just in the middle of the author list. At least in NLP, which can be seen in the generally high number of authors in papers.
rcxdude 2 hours ago [-]
The standards also vary by field. It's not unusual for supervisors to be the last author on any paper that a group publishes, even if they had basically no input into it directly. I've been listed on papers just because I designed and built the equipment that happened to be used for the experiments, even though I didn't do anything but a quick review of the actual paper. For most papers, unless there's an indication otherwise, it's generally only safe to assume the first author listed that actually did the bulk of the work and write-up, and the others are there mainly for having some potentially quite indirect contribution.
nonethewiser 13 hours ago [-]
Isnt this absurd? Say im brainstorming a resume bulet point. Its 15 words. I like it but want to condense it to a single line. I give it to the ai and tell it how much overflows and now it gives me back a simplified sentence and 11 words and some extra stenography constraint? What kind of rule could possibly not effect the quality of that output?
Ok say i do that on 30% of bullet points. Karen the hiring manager is vehemently anti AI. She gives my resume to her AI scanner and what does she find? This not a rhetorical question. Will it treat the text as a whole and not find it? Does it scan every combination of contiguous terms? It could scan bullet points but i could generate in pairs of 2. What about novels?
ruilov 13 hours ago [-]
Suppose I gave you a series of 1000 coin flips. I tell you they were generated by a fair coin. You’re suspicious, you think I hid a watermark in them. But you look at the sequence and 485 are heads, close enough to 50%. You look at the correlation between a coin flip and the next and it’s 0.0433. Close to zero. You do a bunch more stats and everything checks out. So you’re convinced.
Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.
nonethewiser 12 hours ago [-]
But that signal wouldnt carry through if i used your coin for 30% of flips and the results aren’t contiguous right? And for that specific case even 50% and contiguous might not be detected.
It just seems like once you try and detect a signal against something that wasnt one shot it breaks down fast. Unless the signal is constrained to a very small space (pairs of words) but then quality and stealthiness must suffer massively.
Im curious about situations like this.
Paper submitted with some headings, title, and 3 paragraphs. 1 and 3 mostly generated. generated ones have a small percentage of sentences rewritten or deleted. A few find and replaces on terms like load bearing and provenance and glue words around them. The teacher runs the entire thing through a checker. What happens?
Or you have 100 paragraphs and 15 are generated, whole things scanned, what happens?
I guess the implementation and edits matter. And the “resolution “ (ie every sentence-ish bears the mark vs every paragraph). But it seems likely the signal would be lost pretty badly. And i think this is the typical sort of way people actually use AI for important things you might want to validate against.
Absolutely you can hide a watermark in a big chunk of text. But what happens when it’s inside a larger body work that gets checked or split up even a little?
Edit: im reading about synthid and see a splicing doesn’t hurt it much.
Terr_ 13 hours ago [-]
I don't have the math for it, but I wonder how much text/words/tokens are needed before it become easy to embed an ID for authorship as well. There's got to be some number of bytes where it becomes both reliable and hard to find.
"I'm sure glad the LLM was able to fix up my grammar and imagery in the anonymous treatise I made criticizing the authoritarian regime. Hold up, someone's knocking at my door..."
sbszllr 7 hours ago [-]
I've been working in the media and model IP space for quite many years. What this article misses a bit is the threat model for watermarking in general.
Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.
There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.
As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.
EDIT: grammar
rustcleaner 21 hours ago [-]
>watermarks
Whatever happened to just delivering the best product or service? Why must tech be full of ninnying nannies that act against their users, "for their 'safety'‽"
dana-s 7 hours ago [-]
Because the best product should still take care of some social accountability, Meta would not sell their shitty glasses without the light indicator and even with it, as it's trivial to hack there's large social pushback.
jerf 22 hours ago [-]
"What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around."
It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But this scheme could only ever prove that this bit of text was made by a given AI, and validate anything else ever included in the signature hasn't been tampered with. It's not hard to work up a scheme that proves (within reason) a text was generated no earlier than some date by incorporating some sort of information that could only have been known at that date so that could be validated. But this isn't even a step in the direction of proving that something was made by a human. And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
zahlman 22 hours ago [-]
> And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
We don't appear to be talking about cryptographic signing here. (That would never work for the problem because everyone expects unsigned text anyway.) We're talking about:
> It’s basically a text steganography problem (concealing a secret code), made more difficult because the plaintext cannot be arbitrarily manipulated.
As in, trying to force ChatGPT output to contain intentionally crafted ChatGPT-specific LLMisms that a human is unlikely to imitate, even one who reads a lot of ChatGPT output.
baby_souffle 22 hours ago [-]
> It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But how is this implemented? It's a few lines of code to implement a basic "identify and strip/replace any non printable ascii, unicode ..." or whatever.
A screenshot/OCR will also do this.
SO at the end of the day you're left with some dumb rules like "you used `load-bearing` more than once per 500 words, that's AI!"
recursivecaveat 21 hours ago [-]
It's done with a lot more subtlety and embedded directly into the content, no Unicode shenanigans. The basic idea is just to break token generations where the probability is nearly tied in favor of the side that matches the secret key. With a long enough text block you can be statistically certain if the generation was using the key. From another comment: https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...
baby_souffle 14 hours ago [-]
I feel even more vindicated/validated with my general "always proofread/edit the LLM output" policy now :).
jerf 21 hours ago [-]
To quote the original article's context more fully:
"The output of chat tools (and most of the output of AI agents) is not containerized text, but plain old regular text, and so can’t be signed. What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around."
I'm reacting to the idea that "plain old regular text" can't be "signed" because they aren't "files". I'm observing that you can sign a stream of text, and probably other metadata, no problem. To my reading this really is about signing and not stegonographic watermarking, so we're in a context where for some reason the users in question want to carry the certificate of generation by AI and so the fact that this is trivially strippable isn't the issue at hand.
I read it this way because it seems to me clear that it isn't any particularly harder to do the stenographic stuff on a stream than a file (per zahlman's comment), so it only makes sense to be talking about this if we are actually talking about signing.
Waterluvian 14 hours ago [-]
Maybe this is an information theory thing but can there exist a n“ I am human” shibboleth or is this concept fundamentally impossible?
It feel that today you can generally convince someone with text alone that you’re human. Beepity boopity zip zap zoopity today’s AIs aren’t this loose and derpy. Here’s a fTypo and my secret stash of dashes ——-–.
But one day even this won’t do, right?
StevenWaterman 13 hours ago [-]
You'll have to just say something racist, homophobic, anti-Semitic, etc. More intelligence won't "fix" that because the labs don't want to fix that
Waterluvian 13 hours ago [-]
This feels like a Warhammer 40k thing. “We weren’t actually space racists at heart. We just had to say those things to identify against the Necrons. But somewhere along the way the newer generations didn’t see the wink and the nod.”
queenkjuul 2 hours ago [-]
Then we'll just think it's Grok
rsynnott 20 hours ago [-]
“Everyone will simply do fraud” - it is a bizarrely immoral world that the LLMs seem to have unleashed on us. Tech was kinda heading that way anyway, but our friends the magic robots really seem to have turbocharged it.
krapp 19 hours ago [-]
Watermarks have always been a joke, but it is kind of crazy how quickly all of the rules went out the window when LLMs came along. If you'd told someone five years ago that people would just give "AI" root access to their systems and let it run unsupervised except for a list of rules that it sometimes just didn't follow - yeah there's just a nonzero chance that any device with access to your credit card will just buy a thousand eggs. No one knows why, it just happens sometimes even if you beg it not to - you'd be committed.
I remember seeing watermarks in Stable Diffusion ages ago. We're just going to have to come to terms with the reality that AI generated anything will be indistinguishable from human generated anything soon. Counting fingers already doesn't work, most of the "tells" people think work with text are little better than reading tea-leaves. Watermarking anything is futile.
defen 21 hours ago [-]
AI watermarks feel like they're approaching the problem from the wrong side - no matter what it will be possible to remove the watermark (Via manual rewriting, local LLLMs, etc). Instead it seems like we need "proof of human creation". And the only way I see that being possible is hardware-attested proof of keypresses. Which obviously has huge privacy implications, but how else would you actually know that a piece of content was produced by a human pressing keys on a keyboard? The proof would also need to include timestamps for the keypresses, so tell if someone is just copying from another window.
layman51 15 hours ago [-]
That sounds like an interesting idea, but I think there's always weird edge cases. Like what if I pressed the keys to write an essay that was actually generated by an LLM. It would be hardware-attested, but I just transcribed most of my sentences from an LLM and maybe I even introduce my own typos or sentences because I am half-assing the transcription.
baby_souffle 14 hours ago [-]
> Like what if I pressed the keys to write an essay that was actually generated by an LLM. It would be hardware-attested, but I just transcribed most of my sentences from an LLM and maybe I even introduce my own typos or sentences because I am half-assing the transcription.
This is what's known as the "analogue loophole" and it's why DRM and the related "attest, strongly, at every link" schemes are ultimately destined to fail.
Lots of people suggesting we need some sort of cryptographic signature that is unique and unforgeable from every camera; it's the only way to _know_ that something wasn't generated with an LLM!
Just point the camera at a sufficiently high-resolution monitor...
nradov 9 hours ago [-]
There is no need for proof of human creation. That is entirely pointless and irrational. On the Internet, nobody knows you're a dog.
I was assuming it was something like SynthID rather than just sneaky invisible unicode but it's hard to tell from the description.
ch_sm 23 hours ago [-]
here‘s what i don‘t get about this whole discussion. AI companies already store all prompts and responses for future training.
just make an API that returns the string distance between a previously generated paragraph and the query?
that would sidestep this whole problem class.
regulators could even specify how that has to work.
what am i missing?
bonoboTP 21 hours ago [-]
> AI companies already store all prompts and responses for future training.
They claim to only do this when you agree to this in your personal settings. Though Google does say they will train on it, unless you disable history and only use ephemeral chats. Anthropic has a setting for it and claims not to train by default.
Also it would be vulnerable to attacks and privacy problems. You could search for substrings about some suspected information, like "John Smith's medical records show advanced cancer" etc. Of course you'd have to guess the phrasing but still.
ch_sm 2 hours ago [-]
Right, my bad; I know they claim to not store _all_ prompts and responses, that was a bit cynical/hyperbolical of me. However, regulation could force them to do it — with all the downsides that come with that.
Your privacy argument, on the other hand, makes total sense. Not even security measures like homomorphic encryption or locality-sensitive hashing could fix that basic issue.
mirashii 22 hours ago [-]
> AI companies already store all prompts and responses for future training.
They store some prompts and responses, not all, that's what you're missing.
rsynnott 20 hours ago [-]
Well, for a start, they’re probably on dodgy ground there with other European regulations. You’re not really supposed to hold onto hold onto potentially very sensitive data that you have no legitimate interest (a term of art; it doesn’t just mean “I want to keep this”) in indefinitely.
As far as I know, most of these require user consent for retention?
johnnyo 21 hours ago [-]
1. AI company buys and trains on an author’s book when it gets published, it’s now part of the training data.
2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database.
3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.
dpoloncsak 21 hours ago [-]
Ignoring the other technical hurdles of the idea...this problem you outlaid is solved with a timestamp in the database, right? You can easily prove if the prompt was before/after publishing date?
johnnyo 12 hours ago [-]
If it’s before publication date, does that prove the author used AI inappropriately in their writing?
Not really, there are any number of reasons why they might do that.
In my own personal writing, sometimes I run it through LLMs to give grammar or writing suggestions.
echoangle 21 hours ago [-]
Wouldn't this become pretty difficult to do at scale over time? Is there a way to compare the similarity in a database of responses without doing a search over every entry and comparing them? Because that would probably become pretty slow if literally every LLM output is saved and has to be scanned.
dpoloncsak 23 hours ago [-]
Local models?
dTal 23 hours ago [-]
Local models enable:
- watermark-free generation
- the stripping of watermarking from the output of SAAS models
Any discussion of watermarking is dead in the water in a world where we are permitted to have these things. I fear for the future.
graemep 2 minutes ago [-]
Very few people will use local models. In practical terms most people use a few big providers.
dpoloncsak 21 hours ago [-]
Local models also enable:
- not being locked into a provider
- not being forced to have your prompts saved by a possible competitor
- an alternative to the duolopy we quickly see forming
- offline access
I fear for the future without local models, much more than the future with them, and would rather everyone had access to a local model than be certain we catch everyone copy+pasting LLM responses. Watermarking would be cool, but it's not worth losing local for.
dTal 21 hours ago [-]
To be clear, we agree. The problem is that unless local AIs become "normie friendly" real damn quick, we're gonna lose em, because they're damned inconvenient to power. That is what I fear.
dabbz 14 hours ago [-]
I was going to interject that very sentiment. The upfront costs of running any model of value to the average consumer is so large to be basically non-existent. We need hardware to become commoditized and capable of handling LLMs at the scale necessary to enjoy these agentic workflows.
dTal 13 hours ago [-]
Mmm, software too though. Even those with capable hardware (Macbook Pro for instance) are usually incapable of getting over the hump of assembling an inference framework, agentic harness etc. If that ever becomes easy, watch out! Governments will not appreciate the user/developer dichotomy being blown away.
dpoloncsak 20 hours ago [-]
My apologies, I misunderstood your original comment to mean "We shouldn't have local models if it breaks watermarking" but yes, I think we're on the same page now
bethekidyouwant 14 hours ago [-]
I don’t get how you’re gonna lose them how is a guy from Brussels gonna come to your house and turn off your computer?
dTal 13 hours ago [-]
In the short term, governments can suppress free distribution of open weight LLMs. For instance, they could "ask" Hugginface to require an account with a university affiliated email address to download anything. It won't stop you outright, but it will force it underground, slowing everything to a crawl.
In the long run... well, general purpose computing is under threat anyway. Platforms with locked bootloaders and mandatory digital signatures outnumber those that don't. Even on "PCs", the openness is merely a cultural norm observed for a particular market segment[0] by Apple and Microsoft - for they are the only ones with true root keys. Even a brand new motherboard comes with Microsoft signing keys pre-flashed, and you are permitted to enroll your own only by grace. The technical infrastructure to flip a switch and lock it all down is now in place, ready to go at the stroke of a pen. Will you be able to get Llama.cpp from "the app store"? I hazard not. If you think this sounds hyperbolic, look at what is happening to Android.
[0] They don't even segment it the same! Apple segments by "does it have a keyboard", and Microsoft segments it by "does it have a native x86 processor". Apple runs different software on the same basic hardware, Microsoft runs the same basic software on different hardware, but both have created artificially restricted second class computing categories.
graemep 26 seconds ago [-]
They also do not really care if a few geeks get around the system.
bethekidyouwant 9 hours ago [-]
Well I guess all your points are true for Europe but I don’t see how for America. (because freedom) it will certainly cost more but also it already does.
dTal 2 hours ago [-]
Hah. Apple, Microsoft, and Google are all American companies.
nradov 9 hours ago [-]
No need to fear, the future looks bright!
thisoneworks 22 hours ago [-]
Chill my dude. This is just a sane default which will catch normies copy pasting stuff from claude and chatgpt. It's good enough.
dTal 21 hours ago [-]
Do you not want "normies" to have local AI? Do you believe that your ability to use esoteric software makes you special, makes you immune to law? The days where you can find refuge in running "weird hacker crap" instead of 100% goobermint approved software are numbered because news flash: AI means there are no normies anymore. Anyone can do anything.
No, I will not chill. The war on general purpose computing is gonna get real hot real soon.
bethekidyouwant 14 hours ago [-]
Pasting it where?
queenkjuul 1 hours ago [-]
Homework assignments, spam emails, LinkedIn
tzury 21 hours ago [-]
If everyone is using ai (per reported revenues) and nobody like the outcome (banned here, dismissed there), then what is the future of AI?
Gigachad 14 hours ago [-]
The internet will become entirely AI bots talking to each other with no humans writing or reading.
krapp 19 hours ago [-]
That is the future of AI. Everyone will use it whether they like it or not because there won't be an alternative.
phendrenad2 15 hours ago [-]
I have a feeling this is really aimed at coding comments. I've noticed that Opus and Fable have started writing huge block comments all of a sudden.
marginalia_nu 4 hours ago [-]
There's been a notable and well noted shift in Claude's vocabulary toward words that aren't quite appropriate, sometimes completely invented. Biggest claude smell at this moment is that it often uses these weird synonyms from fields like law and medicine in places where you'd expect highly specific well defined words in a technical text.
From a SynthID (or analogous) perspective, it makes a lot of sense you'd see something like this especially in something like a code comment where there's a finite amount of things to say, and it should be said in a terse manner using well defined language. The room for poetic flourish is very small, so you get... 'substrate' or 'verdict' or 'firmament' or whatever else it's using these days.
queenkjuul 1 hours ago [-]
Not only are they huge, they often don't even really make sense, or assume the reader has all the same context as the person who wrote the code to begin with.
Basically 90℅ of my PR review comments now are "reword or delete this comment block, it's not helpful as-is"
fwip 20 hours ago [-]
I don't think this is a good argument. There exist digital watermarking techniques for image and video that are imperceptible to humans but still survive cropping, rotation, resizing, recompression, or an analog round trip (photographing the image or pointing a camera at the video). These watermarks aren't, like, hiding in the low-bits of color information, they're spread among many perceptible details.
It's not clear to me that it's impossible, or even especially difficult, to make something that survives a casual LLM paraphrase. Remember, all you need to encode is a single bit of info. There's a lot of space to redundantly encode that signal.
mmooss 22 hours ago [-]
For SynthID and similar solutions, there is much I don't understand ...
Here's what I grasp: The AI system scores each token and then selects tokens based on those scores. If we encode something in the token selection routine ('in order choose the 1st, 3rd, 1st, 5th, 2nd, then 1st highest scored tokens'), we can identify AI-generated text by comparing sample text (ST) to the expected text (ET) for that prompt.
1) How do we score the tokens for the ET without the original prompt? Even a Markov-like process needs to start somewhere.
2) To recreate ET don't we need to maintain, until the end of time, the AI state - entire model and code - at the time of ST output?
3) Doesn't #2 require maintaining all states for all AIs? Often you won't know when and from which AI system the ST might have been generated. What happens when an AI vendor goes out of business?
4) To recreate ET, don't we effectively have to rerun the prompt? Won't rerunning it for every verification increase most costs of AI output by an order of magnitude? Most of what AI vendors do would be ST validation.
pluto_modadic 21 hours ago [-]
1. score by the past paragraph (this is also how the verifier checks without needing a GPU, it takes the previous length of text)
2. same answer
3. no, it just needs the previous text, private key, and the matrix math (CPU is fine)
4. no, see above.
mmooss 21 hours ago [-]
Thanks ...
1. So we can't score the first paragraph (or similar-sized block), and not short texts? Not deal-breaker, but a limitation.
2. Doesn't the score vary by each AI system state - its model, programming, harness, etc.? Claude's output today doesn't match Gemini's, nor Claude from 2 years ago.
charlieyu1 22 hours ago [-]
Good. Tracking and surveillance have no place in the modern world.
deadbabe 22 hours ago [-]
Society can overcome this problem by changing the way we think about text. Raw plain text should be banned, all text is cryptographically signed by the editor, gui element, or tool that created it.
bonoboTP 22 hours ago [-]
You can ask an AI agent to type into Microsoft Word via computer use.
TheCoelacanth 18 hours ago [-]
That sounds dystopian as fuck.
aaron695 13 hours ago [-]
[dead]
ramesh31 23 hours ago [-]
Yeah but it's like saying "Masterlocks will always be easy to pop off with a hammer". Of course, but by doing so you are actively engaging in fraud, which then puts the onus on you and whoever you are attempting to deceive.
buf 23 hours ago [-]
Except this isn't like saying that at all. This isn't fraud, because it's legal in almost every circumstance.
ramesh31 23 hours ago [-]
>almost every circumstance
Key phrase. And I'm not saying fraud in the legal liability sense. If you're not trying to hide the fact that something was LLM generated, then you have no reason to remove it. If you are trying to hide it, then there's probably a reason, i.e. you would face consequences for doing so, therefore it is fraud.
burnte 23 hours ago [-]
Or, you simply edit the text the LLM generated ruining the hidden message. That's not even remotely fraud. There are legitimate reasons to edit text. There are fewer legit reasons to bash off someone else's lock with a hammer.
Currently I can recognize AI text because I read thousands of ai generated text. I know that 110% of yahoo finance news is generated. I don't want to read an AI generated personal blog, but if I do what's the problem really? Other than the companies distinguishing AI text for getting better training data, how do people benefit from watermarked text exactly?
Really, Anthropic (and the other labs that follow) are just trying to satisfy the requirements of the law so they can continue to serve the EU. Wether it's actually effective is something else entirely.
Malicious compliance is to make cookie acceptance much easier than refusal. Lack of oversight doesn’t punish this. Now, the incentive is to make life hard for people who decline cookies. That results in more annoying pop-ups.
This will most likely bring a sift end to at least the low hanging fruit.
The only way to prevent AI cheating is to make the assignments in class in pen and paper. To be honest, I don't understand why schools moved away from that in the first place.
Cost and convenience.
Being able to work on computers does help people with bad handwriting, which includes some disabilities.
And taking away writing from people with bad handwriting (if not related to disability) is just about the worst thing to do if you want them to ever get better at it.
How does good handwriting improve people's live significantly? I know some people who prefer handwritten notes and I think there is research that says they can be more effective when learning, but most adults very rarely write more than a few lines by hand.
Disabilities don't disappear after school, and children who struggle with motor control in school will continue struggling with it on adulthood. Half the point of school is finding out your weaknesses and learning to overcome them in a safe and controlled environment.
Pushing someone with a disability to improve their handwriting will not cure their disability. How does making people struggle more help them? its like suggesting kids in wheelchairs should be made to try to walk.
Companies will implement it in the worst way possible? Nowhere in the cookie laws does it say you need to add a banner, you just can't spy on users without their consent.
Had google offered a privacy-preserving ad model (basing the ads off the content of the page instead of who the user is) things would be very different.
Does it sound good?
Either anthropic does not tell anyone the signal and only they can verify the text. Or they share the signal and everyone can verify it, which means everyone can bypass it and its just an inconvenience.
The people that make the laws and the people complaining about AI don't know or care about the reality of the situation nearly as much as they care about being re-elected and feeling good about their social posture.
Well, the negative effects of AI are already visible, and I'm not talking about job losses here. I'm talking about AI hallucination, about mass generated spam, about AI-assisted scamming (apparently Indian scammers are shifting to use AI systems [1]) and about political manipulation.
Governments usually are slow to react, but eventually they do react, when a technology poses more issues than it is worth.
[1] https://timesofindia.indiatimes.com/gadgets-news/new-ai-scam...
Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
How do you think you will know the text was AI generated but the person trying to deceive you won’t be able to undo the watermark?
This article says any AI watermarking can be defeated trivially, by running the text through a tool which strips weird Unicode characters (fine) and then runs it through an LLM which replaces all words with similar words (whaaattt?). Running text through another - probably much lower quality llm which scrambles the words you use would make slop even sloppier. And it’s another whole step you have to know to do. A great many people who use LLMs to avoid doing work will not know about these extra steps, or not make use of these sort of tools.
I would say that's not necessarily the case unless there are zero false positives. In fact, your university situation is exactly where a detection system that works 50-90% of the time would be a nightmare if a meaningful share of the 10-50% errors were false positives.
Im very confused how this is even supposed to work at face value.
1) If the verification can be done by anyone, then anyone can bypass it.
2) If it can only be done by anthropic then the government or whomever has the special privilege (not everyone otherwise this is just #1) has to make a specific request
Is the point of the legislation to accurately classify text in general or to simply detect the true positives?
EDIT: Sorry, I misread your comment above. Yeah, hopefully orders of magnitude fewer kids than the number who are getting falsely accused of cheating with LLMs now.
Increasing the accuracy of these systems - both in terms of false positives and false negatives - seems like a good thing.
A proper watermarking system should be able to have an arbitrarily high accuracy - as many nines as you want. And it should be able to actually report the accuracy of its judgements.
If you're worried about kids being falsely accused of cheating using LLMs, you should be cheering on these developments.
Seems like about 3 collisions per 100 million pictures. If everyone have 1000 pictures that is 3 collisions per 100 000 users.
Raiding 2 people in my city for made up CSAM pictures would be way to high false positive rate.
But you can also make any picture match a CSAM hash by adding picked noise.
Probability of a false positive is 3 x 10^-5 in one example.
* https://imgur.com/h5auUhG.jpg
* https://arxiv.org/abs/2301.10226
If some kids are false positives (detected as using ai but didn’t)then how are they lazy?
What about false positives? Imagine being a honest student and then the software declares your work to be AI-generated. How do you defend against that claim? The detection software is a black box and is likely running as a cloud service, so that you have no realistic options for reverse-engineering the false positive detection.
Nothing will save us. You can't automate trust.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
For Anthropic, it’s highly likely that the watermark is a SynthID type mark similar to the one that Sean is talking about (I actually ran the analysis here https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...). When we get confirmation of whether all models actually are watermarked, I think we’ll be even more confident.
Of course it’s always possible that Anthropic has come up with a proprietary scheme, but I think it’s definitely harder to implement.
I think the game will be a cat and mouse game similar to LinkedIn and other websites trying to block scrapers: each iteration makes it harder for someone to figure out the watermarking scheme, but likely not impossible
https://www.youtube.com/watch?v=Iw3G80bplTg
Serious question… what is the point if you cant verify the watermark? Then only the model provider will know.
Is that what this law is about? I thought it was so anyone would know (which of course means anyone can bypass).
Doesnt this and probably all techniques require the validator to know which portion of the text to validate?
If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm?
In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
Catch true positives? Sure
Reliably determine if text was generated by AI? Not even a little.
There's also literal language translation. Generate in language A, translate to language B (Either "manually" by being proficient in it, or with non-LLM translation).
A very, very significant portion of the world knows more than 1 language.
And the more hoops you make cheaters jump through, the better.
Perhaps, but no tools needed for those who know more than one language, which is a lot of people.
Watermarking will never be a perfect tool. But there’s a lot of value in making low effort llm slop detectable. Even if high effort llm slop is still undetectable. Don’t make perfect the enemy of good.
Even if you find a way to 100% watermark any text, couldn't you just use a non-watermark model? I have a hard time believing every AI company on heart would comply.
Hell, even if every AI company on earth decide to somehow apply watermarking to their next model, they would also need to apply it to all the previous version that they commercialize. Given that anybody could make a copy of an open source model right now and would be safe forever, this seems quite the lost cause to me.
The students who, without a single neuron activating, will just copy paste questions to an AI web UI and paste the answers back.
Or the mid-tier corporate people who file 50 page "reports" that are 100% AI bullshit.
None of them will install a local model with no watermarking or bother looking up some Badonk AI service to do the same thing.
People underestimate the value of rules that only take malice and a little knowledge to break.
And they tend to exaggerate that underestimation if they... don't like the rule.
Most people are generally lazy. They upload text to LinkedIn full of "genuine", "honest" and "load-bearing".
And the people who get LLMs to make content for them are often the laziest of us all! Some people can’t even be bothered to remove their prompt from the output. Or they submit work which ends with “is there anything else you want to know about ___?”
Go ahead. Make tools to rip off the llm stenography. Many of the people who would proactively use tools like that are writing content by hand anyway.
Yes, this watermark will be easy to strip. It is still valuable for the vast majority of times where people just don't.
Of course I would want my code formatting tool to normalize that all to plain 0x20 spaces. But it would still be a helpful "brown M&M" test of did you even read CONTRIBUTING and run the code formatter before submitting this PR?
E.g. research paper, law makers, lawyers, state policies, notaries,...
These are much longer content and thus statistically they will disclose a better guess at AI generated content.
Asking another AI to paraphrase will not erase the mark (which they are unaware about) but rather cumulatively add their own mark and make it easier to detect.
The problem is not to use AI, but to endorse the responsibility of the content you (as a human) deliver and somehow make sure that fake-news, biased content or unverified output is detected as early as possible.
Fictive example: imagine the bias is to inject notion of colorfullness in the text regardless or its content:
"I eat an apple" becomes "I eat a red apple"
The 2nd AI is about injecting size component, so the text becomes " I eat a big red apple"
There you get both watermarks. It is not that obvious obviously
Maybe some part of it is that the deluge of slop is uncovering how poorly/sloppily these social institutions were working in the first place.
It is as always (think cybersecurity) the cat and mouse game, what is a weapon is also a defense. You, the simple fact that you are on this very site, means that you are probably more educated to AI than most, so it is not necessarily you that will really benefit from any constraining framework.
There are many people believing in many weird theories and those are more prone to be convinced by a nice narrative. AI did not introduce that, it just made it easier and available to anyone with any intent.
It makes it harder, because amount of bullshit goes up and it is easier to generate plausibly sounding bullshitm
This wouldn't be helped by an AI watermark. What would help is if the reviewer used AI to look up the authors and see the authorship claims are dubious. The papers are still up.
I do think the academic publishing field is corrupt, which is why I'm not convinced it should be on AI providers to help bail it out of doing its one job (verification and trust).
I'd argue all of these could justify authorship, even if they're just in the middle of the author list. At least in NLP, which can be seen in the generally high number of authors in papers.
Ok say i do that on 30% of bullet points. Karen the hiring manager is vehemently anti AI. She gives my resume to her AI scanner and what does she find? This not a rhetorical question. Will it treat the text as a whole and not find it? Does it scan every combination of contiguous terms? It could scan bullet points but i could generate in pairs of 2. What about novels?
Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.
It just seems like once you try and detect a signal against something that wasnt one shot it breaks down fast. Unless the signal is constrained to a very small space (pairs of words) but then quality and stealthiness must suffer massively.
Im curious about situations like this.
Paper submitted with some headings, title, and 3 paragraphs. 1 and 3 mostly generated. generated ones have a small percentage of sentences rewritten or deleted. A few find and replaces on terms like load bearing and provenance and glue words around them. The teacher runs the entire thing through a checker. What happens?
Or you have 100 paragraphs and 15 are generated, whole things scanned, what happens?
I guess the implementation and edits matter. And the “resolution “ (ie every sentence-ish bears the mark vs every paragraph). But it seems likely the signal would be lost pretty badly. And i think this is the typical sort of way people actually use AI for important things you might want to validate against.
Absolutely you can hide a watermark in a big chunk of text. But what happens when it’s inside a larger body work that gets checked or split up even a little?
Edit: im reading about synthid and see a splicing doesn’t hurt it much.
"I'm sure glad the LLM was able to fix up my grammar and imagery in the anonymous treatise I made criticizing the authoritarian regime. Hold up, someone's knocking at my door..."
Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.
There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.
As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.
EDIT: grammar
Whatever happened to just delivering the best product or service? Why must tech be full of ninnying nannies that act against their users, "for their 'safety'‽"
It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But this scheme could only ever prove that this bit of text was made by a given AI, and validate anything else ever included in the signature hasn't been tampered with. It's not hard to work up a scheme that proves (within reason) a text was generated no earlier than some date by incorporating some sort of information that could only have been known at that date so that could be validated. But this isn't even a step in the direction of proving that something was made by a human. And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
We don't appear to be talking about cryptographic signing here. (That would never work for the problem because everyone expects unsigned text anyway.) We're talking about:
> It’s basically a text steganography problem (concealing a secret code), made more difficult because the plaintext cannot be arbitrarily manipulated.
As in, trying to force ChatGPT output to contain intentionally crafted ChatGPT-specific LLMisms that a human is unlikely to imitate, even one who reads a lot of ChatGPT output.
But how is this implemented? It's a few lines of code to implement a basic "identify and strip/replace any non printable ascii, unicode ..." or whatever. A screenshot/OCR will also do this.
SO at the end of the day you're left with some dumb rules like "you used `load-bearing` more than once per 500 words, that's AI!"
"The output of chat tools (and most of the output of AI agents) is not containerized text, but plain old regular text, and so can’t be signed. What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around."
I'm reacting to the idea that "plain old regular text" can't be "signed" because they aren't "files". I'm observing that you can sign a stream of text, and probably other metadata, no problem. To my reading this really is about signing and not stegonographic watermarking, so we're in a context where for some reason the users in question want to carry the certificate of generation by AI and so the fact that this is trivially strippable isn't the issue at hand.
I read it this way because it seems to me clear that it isn't any particularly harder to do the stenographic stuff on a stream than a file (per zahlman's comment), so it only makes sense to be talking about this if we are actually talking about signing.
It feel that today you can generally convince someone with text alone that you’re human. Beepity boopity zip zap zoopity today’s AIs aren’t this loose and derpy. Here’s a fTypo and my secret stash of dashes ——-–.
But one day even this won’t do, right?
I remember seeing watermarks in Stable Diffusion ages ago. We're just going to have to come to terms with the reality that AI generated anything will be indistinguishable from human generated anything soon. Counting fingers already doesn't work, most of the "tells" people think work with text are little better than reading tea-leaves. Watermarking anything is futile.
This is what's known as the "analogue loophole" and it's why DRM and the related "attest, strongly, at every link" schemes are ultimately destined to fail.
Lots of people suggesting we need some sort of cryptographic signature that is unique and unforgeable from every camera; it's the only way to _know_ that something wasn't generated with an LLM!
Just point the camera at a sufficiently high-resolution monitor...
I was assuming it was something like SynthID rather than just sneaky invisible unicode but it's hard to tell from the description.
just make an API that returns the string distance between a previously generated paragraph and the query?
that would sidestep this whole problem class.
regulators could even specify how that has to work.
what am i missing?
They claim to only do this when you agree to this in your personal settings. Though Google does say they will train on it, unless you disable history and only use ephemeral chats. Anthropic has a setting for it and claims not to train by default.
Also it would be vulnerable to attacks and privacy problems. You could search for substrings about some suspected information, like "John Smith's medical records show advanced cancer" etc. Of course you'd have to guess the phrasing but still.
Your privacy argument, on the other hand, makes total sense. Not even security measures like homomorphic encryption or locality-sensitive hashing could fix that basic issue.
They store some prompts and responses, not all, that's what you're missing.
As far as I know, most of these require user consent for retention?
2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database.
3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.
Not really, there are any number of reasons why they might do that.
In my own personal writing, sometimes I run it through LLMs to give grammar or writing suggestions.
- watermark-free generation
- the stripping of watermarking from the output of SAAS models
Any discussion of watermarking is dead in the water in a world where we are permitted to have these things. I fear for the future.
- not being locked into a provider
- not being forced to have your prompts saved by a possible competitor
- an alternative to the duolopy we quickly see forming
- offline access
I fear for the future without local models, much more than the future with them, and would rather everyone had access to a local model than be certain we catch everyone copy+pasting LLM responses. Watermarking would be cool, but it's not worth losing local for.
In the long run... well, general purpose computing is under threat anyway. Platforms with locked bootloaders and mandatory digital signatures outnumber those that don't. Even on "PCs", the openness is merely a cultural norm observed for a particular market segment[0] by Apple and Microsoft - for they are the only ones with true root keys. Even a brand new motherboard comes with Microsoft signing keys pre-flashed, and you are permitted to enroll your own only by grace. The technical infrastructure to flip a switch and lock it all down is now in place, ready to go at the stroke of a pen. Will you be able to get Llama.cpp from "the app store"? I hazard not. If you think this sounds hyperbolic, look at what is happening to Android.
[0] They don't even segment it the same! Apple segments by "does it have a keyboard", and Microsoft segments it by "does it have a native x86 processor". Apple runs different software on the same basic hardware, Microsoft runs the same basic software on different hardware, but both have created artificially restricted second class computing categories.
No, I will not chill. The war on general purpose computing is gonna get real hot real soon.
From a SynthID (or analogous) perspective, it makes a lot of sense you'd see something like this especially in something like a code comment where there's a finite amount of things to say, and it should be said in a terse manner using well defined language. The room for poetic flourish is very small, so you get... 'substrate' or 'verdict' or 'firmament' or whatever else it's using these days.
Basically 90℅ of my PR review comments now are "reword or delete this comment block, it's not helpful as-is"
It's not clear to me that it's impossible, or even especially difficult, to make something that survives a casual LLM paraphrase. Remember, all you need to encode is a single bit of info. There's a lot of space to redundantly encode that signal.
Here's what I grasp: The AI system scores each token and then selects tokens based on those scores. If we encode something in the token selection routine ('in order choose the 1st, 3rd, 1st, 5th, 2nd, then 1st highest scored tokens'), we can identify AI-generated text by comparing sample text (ST) to the expected text (ET) for that prompt.
1) How do we score the tokens for the ET without the original prompt? Even a Markov-like process needs to start somewhere.
2) To recreate ET don't we need to maintain, until the end of time, the AI state - entire model and code - at the time of ST output?
3) Doesn't #2 require maintaining all states for all AIs? Often you won't know when and from which AI system the ST might have been generated. What happens when an AI vendor goes out of business?
4) To recreate ET, don't we effectively have to rerun the prompt? Won't rerunning it for every verification increase most costs of AI output by an order of magnitude? Most of what AI vendors do would be ST validation.
2. same answer
3. no, it just needs the previous text, private key, and the matrix math (CPU is fine)
4. no, see above.
1. So we can't score the first paragraph (or similar-sized block), and not short texts? Not deal-breaker, but a limitation.
2. Doesn't the score vary by each AI system state - its model, programming, harness, etc.? Claude's output today doesn't match Gemini's, nor Claude from 2 years ago.
Key phrase. And I'm not saying fraud in the legal liability sense. If you're not trying to hide the fact that something was LLM generated, then you have no reason to remove it. If you are trying to hide it, then there's probably a reason, i.e. you would face consequences for doing so, therefore it is fraud.