>she wrote she had simply asked ChatGPT to make “a New Yorker-style cartoon.”
A "style" can't be copyrighted, at least in US law. They might have a stronger case of trademark/likeness infringement, but the fact that the person knew it was AI generated would make that difficult. Of course they knew it wasn't made by Brendan Loper. Of course, if they then published it, the other people viewing it might not know this, but who published it?
Nobody seems to care about AI automating coders out of a job.
If anyone can just prompt all their basic “information needs” however how sloppy, then what remains of the economy? Health care, child care, handyman?
Most people won’t even pay for ad free YouTube. I don’t think any software business can survive AI as a substitute good even if it’s inferior (and it might not be).
You may not like AI but their business model is a silly twitter-worthy dunk. This is obviously not their business model. I didn't know I could've plagiarized all this code myself the entire time!
This is one of the reasons I'm extremely unimpressed by complaints from openai and anthropic that other labs are "distilling" their models based on them... Basically: "You're training your model by running it again our own model which is itself a gargantuan copyright and content violation of a scale never seen before by humankind, HOW DARE YOU"
>I'm certain if I drew a cartoon and used the signature that looks just like a real cartoonists', I would get sued and liable for the damage.
Would you? Always? Suppose I hate Obama drone striking people, so I made a satirical cartoon of him signing an executive order to "bomb brown people" or whatever, affixing his signature[1] to that image. Would that get me in trouble, even if the image was clearly satirical? What if someone takes that, then passes it as non-satire?
This has been a perennial problem with my own generated comics with both Nano Banana Pro and ChatGPT (all generations). I often have to put in an extra edit to erase the false signature. It is annoying and I'm unsurprised most users don't bother.
Despite what all the clickwrap warnings and "AI can make mistakes" subtitles might lead you to believe, the service offering of AI is explicitly designed to be as "one and done" as possible. The inherent nature of these tools is to service laziness, and disincentivize too much scrutiny.
And then people wonder why the default mood of AI is so pessimistic. It's just revealing all of society's broken windows and adding a few more in the process.
This reminds me my early attempts to use GitHub Copilot when it just straight added some guy's name in a javadoc copyright note in the code it generated.
Who knows what happened, kids sometimes become suicidal... and 1% of the population that are psychopaths are on occasion documented killing people over $5 because they uniquely don't value most living things except themselves.
What is weird was how fast things were buried, and the family's concerns were never properly addressed. =3
> The New York Times article cites Stanford University law professor Mark Lemley, who disagreed that generative AI services violate copyright law, and intellectual property attorney Bradley Hulbert, who said a new law might be necessary to settle the question of legality.
> Months after Balaji's death, which attracted significant public attention, Hulbert told Fortune magazine that Balaji's essay "[reads like] the argument of a really smart non-lawyer who read up on the subject but does not have a thorough understanding".
If there's some kind of industrial-scale intimidation campaign that's stopping IP lawyers from litigating the case of their lifetime, that's an even bigger story than OpenAI taking out a hit on somebody. It seems like they're agreeing that the copyright abuse was never hidden, and it's sufficiently transformative enough that nobody could argue it's illegal.
That seems to suggest that Disney's lawyers agree. You can use AI to violate copyright laws no different from a text editor or Bittorrent, but training it on copyright material isn't inherently illegal.
Engineering manager at my company put a comic at the end of our sprint demo that was signed bloper. Except it wasn’t funny at all, and kind of weird. I asked him, sure enough it was ChatGPT and he didn’t notice the signature.
Nice. This demolishes the "LLMs can reason" (but not enough to avoid this sort of basic error) and "humans make mistakes too" (not like this) talking points from the LLM promoters.
The thing about a generative language model that’s trained from a massive but unknown corpus is, it’s practically (if not theoretically) impossible to evaluate the extent to which data leakage contributes to any particular output.
But I would argue that, as things currently stand, “sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus” remains a more parsimonious explanation than “it’s doing actual reasoning” for how this neural network architecture produces the phenomena we’ve been observing.
This is completely silly. If you don’t think LLMs can reason, you’ve either never used them to do tasks that require reasoning, or you don’t understand enough to recognize what’s involved in the responses you get.
In this case it’s clearly the latter, because you’re confusing image generation models with LLMs. There are very big differences between the two. No-one is claiming that image generation models are capable of reasoning.
They are not reasoning, stop referring to it as behaving like a human. It does nothing of the sort. FFS lmao.
An airplane does not flap wings but flies. Know the difference. In many respects humans do not care about 1-to-1 mapping of the production process but the output.
What's the argument here? An airplane, a bee and a bird all fly despite doing it totally differently. LLMs also reason despite being made out of matmuls instead of meat.
The most common argument for this is some core unexamined axiom that only humans can reason by definition, and then working backwards to a justification for that.
What do you think they are doing, and do you think machines can reason (in general, not necessarily current systems)? If they can't, how do you explain humans being able to reason given that we are physical machines too?
> “It’s like somebody attributed a quote to me that I didn’t say.”
> Katzenstein considers the reproduction of his signature by ChatGPT to be more than just a violation of intellectual property; to him, it’s closer to false impersonation. “[ChatGPT] is attaching my name to work that I do not endorse or like. It’s slop, and unlike the other slop that I’ve encountered, this is slop that’s pretending to be me.”
> “I’ve had people hack my credit card,” said Joe Dator, a New Yorker contributor for the past 20 years. “That feels like less of a violation than this. When they hacked my credit card, they didn’t dress up like me.”
So this has morphed from plagiarism and copyright infringement (bad) to impersonation (also bad, arguably worse, and maybe more provable in court). It’s chilling to think of the implications of having one’s signature attached to a document or to words that are not one’s own.
And I think maybe it's time for that. People need to learn that there's real, expensive legal liability for doing stuff like this. And AI companies the same.
I am very much not an advocate of "sue everybody for everything". This is major enough that it clears my threshold.
I've seen photorealistic generations add (distorted but still somewhat recognisable) watermarks too, because that's what the training data had.
Everything is a derivative work, and always has been. AI is just making that salient fact so much more visible, and now everyone who believes in the delusion of Imaginary Property is scared at that truth revealing itself.
Incidentally, this is also what young humans learning to draw will do. They start by copying what they've seen.
Whatever the original intention, this is clearly a bug and should be fixed. But should ChatGPT sign its cartoons with its own name or leave them unsigned?
I think what you're seeing is the probability of a particular signature or style of signature appearing on a particular style of cartoon, not an intent to sign.
This. It's also while you'll sometimes get a mangled Getty Images watermark on some image generations, or a logo in the bottom left corner. If it's a prominent feature in the training dataset it'll show up, exactly how these models are supposed to work.
The 'bug' here is whatever post-processing step or system prompt is in place to steer the model away from doing this.
No one should be allowed to claim they drew something when they didn’t draw any part of it and LLM’s are not people/can’t work without a person. We don’t credit pens and paintbrushes after all.
One could argue nobody should be allowed to claim it. It just exists.
Even if the person can't claim the copyright of the image produced they ARE responsible for the use of their tools and what they do with the output.
In this case, they released an image with someone else's signature on it. That is wrong, the person should take the blame for that.
The person releasing the image may take it up with the AI service that their tooling led them into making such a mistake. But good luck with that in court...
Why is it a bug? If other parts of the generated illustration are similarly taken from an artist, why not the signature as well? Why is a signature crossing the line but the rest of the image isn't?
For the same reason I'm allowed to draw, paint, or write things very similar to what others have drawn, painted, or written but I have to sign my own name not theirs.
You are a person, LLMs are not. You know this, which is why you know that if you signed someone else's name it would be forgery, but when you see the machine do it you call it a bug.
If the machine is like you, the machine is a forger. The machine is not like you, it is simply blending the work of others to order. Adding someone else's signature is simply part of that statistical process.
It seems like all the criticisms of Gen AI and LLM seems to concentrate on OpenAI and their products over products from anthropic and others.
I am pretty sure that this faking of signature can be done by Gemini, claude as easily as chatgpt.
I follow anti-LLM discourse quite a lot, and across the main bulletpoints: energy/carbon emissions, content worker harm, job displacement, deskilling, mental health effects, and copyright/plagiarism, the plagiarism one seems to have the most attention, and it's also the most solvable, if there were only more serious effort on ethically sourced models that can actually do the real science / math / code work that is what LLMs are best at. The whole world of LLMs to create videos/books/art/literature is where most of the offense is (the video/imagery side of it is where most of the energy/carbon emissions problems are too. and content worker harm).
I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
Yep. I’d probably be a lot less chastised in some circles for using Claude Code at work if it wasn’t misconstrued as being in support of, I don’t know, encroaching on the hypothetical commissions of a chronically online instagram furry artist or something.
It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.
I am skeptical of there being sufficient data to build “ethical” training datasets, and I’m confident that much of the same contingent will (somewhat rightfully) argue that ‘second-generation’ copyrighted AI material has already irreversibly made its way into every modern dataset.
> It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.
That’s not a justification. If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you “we don’t necessarily disagree with you about polluting the water, but in our field—which you do not understand, and in which the underlying build process is often not the water pollution—what we’re doing presents very real productivity gains”?
> I am skeptical of there being sufficient data to build “ethical” training datasets
Then you don’t build any. What fucked up world we live in where people think it’s OK to be unethical because they want something and can’t think of any other way to do it. What monumentally selfish rotten babies.
I think you can train on math /science using synthetic generation to a significant extent. Training for coding requires more of the "scraping github / stackoverflow" angle but IMO that's a shallower hill to climb than scraping copyrighted art and literature.
There are actual models trained on ethical datasets but they are obviously not very high powered. If companies with the resources of an anthropic or openai were doing it (ha) it would be more feasible
> has already irreversibly made its way into every modern dataset
The "gray goo" scenario finally happens... for AI. That's actually the good ending for humanity. I love it! Poetic and believable. Data doesn't "heal" like nature. :D
2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr.
3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled.
4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas
5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling.
Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech
>“I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains”
Being able to prove such gains in better products would be a start. And an emphasis on how it assists existing engineers/mathmaticians/researchers, not that any accomplishment made with AI assistance is "AI solves problem".
I don't know whatever happened to "words are cheap". I guess it literally made money to say words, so that adage is false for the time being.
>I am skeptical of there being sufficient data to build “ethical” training datasets
Well if all those scam job ads paying 100/hr to create AI training content was not a scam and instead the approach from the start, there may have been a chance to bridge that gap ethically. The industry chose to break things and is trying to act mad that people are mad at all the broken stuff.
These results are entirely a consequences of the actions chosen. And I don't believe there was ever an honest consideration of there being ethical training datasets. They just thought they could brute force society with fearmongering and bribes. The BOTD was already low in the beginning but completely gone now.
I don't think it is likely that they could get enough data without stealing. It would be incredibly costly to have to pay artists to church out art just to train an AI.
I don't really see the difference, code is protected by copyright (or copyleft) as much as art is, and yet the LLM scrapers use it without scruples. Same goes for math and science publications.
In my experience, the science/math/code crowd don't care about copyright/plagiarism as much as the video/art/literature crowd, so the first crowd turns a blind eye to most of the latter crowd talks about.
Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.
While that is true, theoretically a regulation could be enacted that output tokens must focus on STEM research and other practical tasks and the LLM must refuse tasks outside of those areas, just as Claude disallowed cybersecurity tasks. Obviously this would never happen, but the theft of the training data wouldn’t matter as much if the usecases were less sinister.
openai and anthropic trained on actually stolen data since it was pirated datasets.
google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.
i think math is close to art (just to be a contrarian, but kind of really)
its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good
but the art people have no such delusions so they’re just strictly against
imo proof writing is more akin to art than coding/tech but…
> I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.
I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.
This is a pretty solid argument against people who argue that LLMs are more than just (very massive) next token predictors.
If there was any thought or underlying thought going on here not putting a signature (at least a real one) would be the right move, despite it being less likely. It would realize, while generating the pixels that eventually became a signature, that it shouldn't do that.
> This is a pretty solid argument against people who argue that LLMs are more than just (very massive) next token predictors.
This quote is a pretty solid argument that you need to understand the technology you’re trying to criticize better. This issue has nothing to do with LLMs. LLMs are not image generation models.
A LLM, at least, prompted and served this image. LLMs can ingest images. They actually can generate them as well but that probably wasn't done here.
In the course of the conversation a with chatgpt, this image was generated and served by an LLM. It clearly shouldn't have been by any sort of reasoning.
If you do not want to be called a duck, it would help if you stopped quacking like one. Maybe you aren't a duck, but you aren't helping your case with stories like this about how AI generates images.
Article said
>she wrote she had simply asked ChatGPT to make “a New Yorker-style cartoon.”
A "style" can't be copyrighted, at least in US law. They might have a stronger case of trademark/likeness infringement, but the fact that the person knew it was AI generated would make that difficult. Of course they knew it wasn't made by Brendan Loper. Of course, if they then published it, the other people viewing it might not know this, but who published it?
https://commons.wikimedia.org/wiki/Commons:When_to_use_the_P...
We can all try really hard to pretend that's not the business model, but that's totally the business model.
1. https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...
If anyone can just prompt all their basic “information needs” however how sloppy, then what remains of the economy? Health care, child care, handyman?
Most people won’t even pay for ad free YouTube. I don’t think any software business can survive AI as a substitute good even if it’s inferior (and it might not be).
Have people ever really broadly cared about the IT professionals behind their devices?
The same should apply to LLM vendors.
Would you? Always? Suppose I hate Obama drone striking people, so I made a satirical cartoon of him signing an executive order to "bomb brown people" or whatever, affixing his signature[1] to that image. Would that get me in trouble, even if the image was clearly satirical? What if someone takes that, then passes it as non-satire?
[1] https://en.wikipedia.org/wiki/File:Barack_Obama_signature.sv...
For certain values of "my own".
And then people wonder why the default mood of AI is so pessimistic. It's just revealing all of society's broken windows and adding a few more in the process.
What is weird was how fast things were buried, and the family's concerns were never properly addressed. =3
> The New York Times article cites Stanford University law professor Mark Lemley, who disagreed that generative AI services violate copyright law, and intellectual property attorney Bradley Hulbert, who said a new law might be necessary to settle the question of legality.
> Months after Balaji's death, which attracted significant public attention, Hulbert told Fortune magazine that Balaji's essay "[reads like] the argument of a really smart non-lawyer who read up on the subject but does not have a thorough understanding".
If there's some kind of industrial-scale intimidation campaign that's stopping IP lawyers from litigating the case of their lifetime, that's an even bigger story than OpenAI taking out a hit on somebody. It seems like they're agreeing that the copyright abuse was never hidden, and it's sufficiently transformative enough that nobody could argue it's illegal.
Many already settled out of court with Disney due to trademark violations, then killed a popular project mostly used for Star-wars satire at the time.
Best of luck =3
The thing about a generative language model that’s trained from a massive but unknown corpus is, it’s practically (if not theoretically) impossible to evaluate the extent to which data leakage contributes to any particular output.
But I would argue that, as things currently stand, “sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus” remains a more parsimonious explanation than “it’s doing actual reasoning” for how this neural network architecture produces the phenomena we’ve been observing.
This is completely silly. If you don’t think LLMs can reason, you’ve either never used them to do tasks that require reasoning, or you don’t understand enough to recognize what’s involved in the responses you get.
In this case it’s clearly the latter, because you’re confusing image generation models with LLMs. There are very big differences between the two. No-one is claiming that image generation models are capable of reasoning.
They are not reasoning, stop referring to it as behaving like a human. It does nothing of the sort. FFS lmao.
An airplane does not flap wings but flies. Know the difference. In many respects humans do not care about 1-to-1 mapping of the production process but the output.
The most common argument for this is some core unexamined axiom that only humans can reason by definition, and then working backwards to a justification for that.
> Katzenstein considers the reproduction of his signature by ChatGPT to be more than just a violation of intellectual property; to him, it’s closer to false impersonation. “[ChatGPT] is attaching my name to work that I do not endorse or like. It’s slop, and unlike the other slop that I’ve encountered, this is slop that’s pretending to be me.”
> “I’ve had people hack my credit card,” said Joe Dator, a New Yorker contributor for the past 20 years. “That feels like less of a violation than this. When they hacked my credit card, they didn’t dress up like me.”
So this has morphed from plagiarism and copyright infringement (bad) to impersonation (also bad, arguably worse, and maybe more provable in court). It’s chilling to think of the implications of having one’s signature attached to a document or to words that are not one’s own.
And I think maybe it's time for that. People need to learn that there's real, expensive legal liability for doing stuff like this. And AI companies the same.
I am very much not an advocate of "sue everybody for everything". This is major enough that it clears my threshold.
Everything is a derivative work, and always has been. AI is just making that salient fact so much more visible, and now everyone who believes in the delusion of Imaginary Property is scared at that truth revealing itself.
Incidentally, this is also what young humans learning to draw will do. They start by copying what they've seen.
The 'bug' here is whatever post-processing step or system prompt is in place to steer the model away from doing this.
One could argue nobody should be allowed to claim it. It just exists.
Even if the person can't claim the copyright of the image produced they ARE responsible for the use of their tools and what they do with the output.
In this case, they released an image with someone else's signature on it. That is wrong, the person should take the blame for that.
The person releasing the image may take it up with the AI service that their tooling led them into making such a mistake. But good luck with that in court...
If the machine is like you, the machine is a forger. The machine is not like you, it is simply blending the work of others to order. Adding someone else's signature is simply part of that statistical process.
I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.
I am skeptical of there being sufficient data to build “ethical” training datasets, and I’m confident that much of the same contingent will (somewhat rightfully) argue that ‘second-generation’ copyrighted AI material has already irreversibly made its way into every modern dataset.
That’s not a justification. If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you “we don’t necessarily disagree with you about polluting the water, but in our field—which you do not understand, and in which the underlying build process is often not the water pollution—what we’re doing presents very real productivity gains”?
> I am skeptical of there being sufficient data to build “ethical” training datasets
Then you don’t build any. What fucked up world we live in where people think it’s OK to be unethical because they want something and can’t think of any other way to do it. What monumentally selfish rotten babies.
There are actual models trained on ethical datasets but they are obviously not very high powered. If companies with the resources of an anthropic or openai were doing it (ha) it would be more feasible
The "gray goo" scenario finally happens... for AI. That's actually the good ending for humanity. I love it! Poetic and believable. Data doesn't "heal" like nature. :D
But sure, there's um, an ethical way of doing that?
1. Only use open source/CC compliant assets.
2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr.
3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled.
4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas
5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling.
Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech
[0]: https://harpers.org/archive/2025/01/the-ghosts-in-the-machin...
Being able to prove such gains in better products would be a start. And an emphasis on how it assists existing engineers/mathmaticians/researchers, not that any accomplishment made with AI assistance is "AI solves problem".
I don't know whatever happened to "words are cheap". I guess it literally made money to say words, so that adage is false for the time being.
>I am skeptical of there being sufficient data to build “ethical” training datasets
Well if all those scam job ads paying 100/hr to create AI training content was not a scam and instead the approach from the start, there may have been a chance to bridge that gap ethically. The industry chose to break things and is trying to act mad that people are mad at all the broken stuff.
These results are entirely a consequences of the actions chosen. And I don't believe there was ever an honest consideration of there being ethical training datasets. They just thought they could brute force society with fearmongering and bribes. The BOTD was already low in the beginning but completely gone now.
Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.
The current models intelligence depends on massive training dataset of essentially stolen data
google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.
its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good
but the art people have no such delusions so they’re just strictly against
imo proof writing is more akin to art than coding/tech but…
I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.
I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.
If there was any thought or underlying thought going on here not putting a signature (at least a real one) would be the right move, despite it being less likely. It would realize, while generating the pixels that eventually became a signature, that it shouldn't do that.
This quote is a pretty solid argument that you need to understand the technology you’re trying to criticize better. This issue has nothing to do with LLMs. LLMs are not image generation models.
In the course of the conversation a with chatgpt, this image was generated and served by an LLM. It clearly shouldn't have been by any sort of reasoning.
If you do not want to be called a duck, it would help if you stopped quacking like one. Maybe you aren't a duck, but you aren't helping your case with stories like this about how AI generates images.