Pop!_OS bans AI-generated code from much of its codebase

(neowin.net)

113 points | by bundie 22 hours ago

23 comments

  • brink 20 hours ago
    I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
    • lucianmarin 20 hours ago
      Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
      • dnautics 20 hours ago
        I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.

        Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance

      • fasterik 20 hours ago
        I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
        • bitexploder 19 hours ago
          It is also easier than ever to build specs and have nice easy to maintain projects. It just doesn't happen magically via few shot prompts :)
        • globular-toast 19 hours ago
          You could have also copied an SVG renderer that implements the whole spec from whatever open source project the model copied it from.
          • fasterik 18 hours ago
            It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
            • sashank_1509 10 hours ago
              I don’t think we should hold humans and models to the same standards. Humans have a very small working memory. Most humans cannot reproduce code they wrote even a year back exactly.

              A model can reproduce large swaths of its training data exactly. It’s a different algorithm that powers its learning process (it’s why it needs trillions of tokens to even learn basics of language).

              If there was a spectrum from copying on one end to creative production inspired from something else on the other end, the human generally lies heavily on the right end, while the model is much more on the left, that gap is large enough, that yes the model is in some sense “copying”.

      • fignews 20 hours ago
        Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
        • catlifeonmars 19 hours ago
          It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.

          I’m saying it’s probably multiple factors and both you and GP are right.

        • sirsinsalot 19 hours ago
          Have you considered it isn't?

          Save your "you're holding it wrong" if you're not going to suggest how to hold it.

          Cult speak escape hatches are intellectually lazy.

          • fignews 19 hours ago
            Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D

            https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.

            There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.

            I guess they know how to hold it?

            • chmod775 19 hours ago
              You're using quantity metrics to answer a quality question.

              I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.

              If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.

              If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.

            • diek 15 hours ago
              I find it funny that the original complaint was: "AI made a mess of the codebase".

              Your response was: "Well you're not doing it right, but these hermes devs know what they're doing".

              But the blog post you linked to shows their prompt, which is:

              > I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.

              So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.

              And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that's been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.

            • northstar702 16 hours ago
              [flagged]
          • williamcotton 19 hours ago
            I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.

            Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!

            What was your process?

            • hirvi74 16 hours ago
              Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what's there to really talk about when the code can speak for itself?

              In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.

        • runarberg 20 hours ago
          Meaning OP is a better programmer then a statistical model which produces the most probable results?
          • fignews 19 hours ago
            Meaning a bad workman blames his tools.
            • nvme0n1p1 18 hours ago
              A bad workman blames his tools.

              A good workman shuts up and finds better tools without complaining.

              • runarberg 14 hours ago
                An expert workman knows which tools to avoid, and explains it to their peers why to avoid said tools.
      • rickydroll 10 hours ago
        I've had a different experience. My code is better. The ease of refactoring and practicing very defensive programming without getting burned out from the repetitive boilerplate. It all comes together better than it did when I was programming professionally too many moons ago. IFF I don't let the LLM toddler run amok. :-)

        Seriously, I find I need to slow down the rate of change. I don't move forward until I understand the change proposed and have updated the docs. At the same time, I find that keeping up with the LLM/agent is exhausting. 3 hours with an LLM leaves me as tired as 6 hours with a keyboard had previously. I find that coding when tired or fuzzy yields code that shouldn't have been written in the first place. Sadly, once it's been written and debugged, the temporary fix becomes permanent.

      • cyanydeez 3 hours ago
        I've adopted open source projects and have been working exclusively in languages I have very little experience with: typescript and go.

        I'm using local models, and they go slow enough that I have no trouble following along with what they're doing; but visually, both go and typescript, along with react native, make me puke. So I wouldn't be able to do this without AI.

        I describe how to do it in my comment history, but it's basically a Super-TDD along with some custom engineering harness.

        I don't want to say skill issue, but the same way you can give a chain saw to a teenager and one to a skill craftsman, well, AI can obviously create whatever you want it to do.

        I think some of the variety is simply how fast SOTA models pump out garbage that you simply have to close your eyes because it's not sensible to just watch characters flow across the screen.

        Almost all the coding I'm doing via AI is just faster than readable. But I can see the thinking traces and I stop to model when it's obvious it doesn't understand my intent, etc.

        So I'm not doubting you created garbage. I'm just doubting that it's a product of soley AI use.

      • Keyframe 19 hours ago
        I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
      • satvikpendem 20 hours ago
        When and what model?
    • joshheitzman 19 hours ago
      People, especially those who don't do software development, often conflate coding and software development.

      The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.

      That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.

      • PorciiVorbesc 19 hours ago
        >The models aren't good at architecture and design.

        Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?

        Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.

        Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?

        • joshheitzman 17 hours ago
          • PorciiVorbesc 17 hours ago
            Cool. Can you elaborate at which types of task you are better than SOTA LLMs in context of "being good at SW design and architecture"? Got any examples? Is it at the interview questions? Or real world problems? If so what is the scale of the real world SW design and architecture problems you're better than the SOTA LLMs? Is it FAANG scale or mom and pop shop scale?

            And do you consider yourself to be representative of the average developer, above them, or below them?

            LLMs don't even need to be better than the average dev, let alone the top performing ones, like you. If they can just be better than the bottom 20% of devs and white collar workers in general(easily achievable when you've been around the block and saw how many useless people just keep warm chairs for high wages in large companies), that's already a huge win for those products.

            What I mean, at a previous job I had ran into a memory leak issue in our backend and discovered a colleague pushed a library into prod which came with comments in the source code saying "DO NOT USE IN PROD, IT CAUSES A MEMORY LEAK!". There's cases where human stupidity and carelessness far surpasses whatever issues LLMs cause so maybe the average dev isn't really that much better than the SOTA LLMs.

            • joshheitzman 15 hours ago
              I'm currently driving my AI coding agent harness through the process of refactoring itself. This is the last big refactor I need done before I can polish and release it as OSS later this month. So its not a large scale project I'm working on at present. The general problem is that mostly add new code and try to minimize editing existing code, which effectively means the code base is grown organically rather than being intentionally architected and designed. A few of specifics:

              1) duplication - LLMs are great at generating lots of text, so its faster and easier for them to generate entirely new facilities that overlap heavily with existing ones then it is for them find existing facilities that should be expanded and refactored (note I just said 'find'; actually editing raises the time and difficulty even more). This is fine for a while as the duplicate facilities usually work just fine, up until something needs to be changed across all of them and they miss changing one or more of them, things break, and a bunch of tokens have to be burned tracking down the issue.

              2) ever increasing surface area - even when making changes that do expand a facility without much duplication they frequently only add without removing much of anything or changing the overall design of the facility to reduce the amount of state its tracking and the number of branches it has based on that state. I've never seen one decide to split up something large or with too many responsibilities on their own. They will happily create a god class or function and just keep making it bigger.

    • mulemisterX 19 hours ago
      I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
    • Dfiesl 19 hours ago
      Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.
    • satvikpendem 20 hours ago
      As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.

      Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.

    • kees99 6 hours ago
      "Coding is solved" is a very high bar, and current-gen AI/LLM is very far from that.

      Where AI did make a lot of impact is triage. I can throw a messy bug report of an intermittent issue at claude, and tell I need issue reproduced and fix developed, and there is a well above 50% chance it'll deliver. Never commit that fix to the codebase as-is, of course.

    • analog_daddy 20 hours ago
      Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023. I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
      • zanderwohl 19 hours ago
        If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.
    • vjvjvjvjghv 19 hours ago
      I treat AI like an intern or junior team member. With enough guidance, they can contribute a lot but you can't let them loose without supervision or they usually will produce a big mess. As of now, you are still responsible for overall architecture. One strong indicator that something is going wrong are big pull requests where the AI has rewritten large sections of the code.
    • xenadu02 19 hours ago
      AI is a useful tool but it absolutely needs a lot of human guidance.

      Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.

      Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.

    • nicoburns 20 hours ago
      Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
      • thesdev 20 hours ago
        Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
        • satvikpendem 20 hours ago
          It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
        • drewstiff 18 hours ago
          Assuming we are measuring time and

          `total = dev + review`

          If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.

          Will you potentially have a backlog of code waiting for review? Also yes.

          Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.

        • nicoburns 16 hours ago
          Yes, but it's the only option if you want to retain a maintainable code base. And you can always slow down. You'll probably still be a bit faster than you were before.
        • classified 20 hours ago
          Having unmaintainable code faster is only an advantage if it's a one-shot throwaway artifact.
    • TuxSH 18 hours ago
      Vibe coding has always had and always will have one fundamental flaw: you didn't write the code it produced, therefore you don't have a good understanding/mental model of it.

      On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.

      And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.

      Stuff that used to take weeks or months now just takes a few hours, or less.

    • cortesoft 19 hours ago
      I am not saying you are right or wrong, but I don’t think you can make such a broad conclusion simply because of your own experience.

      You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.

      I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.

      However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.

      I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.

      I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”

      I wish people would stop assuming their experience with something is the only possible truth.

      • threethirtytwo 18 hours ago
        Stop being overly polite to these people. It is seriously coming down to the point of utter delusion.

        As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.

    • bitfilped 20 hours ago
      The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
      • orangecat 20 hours ago
        I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
        • mod50ack 20 hours ago
          I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
          • drewstiff 18 hours ago
            I would argue that using an older model and assuming that it is the cutting edge is basically equivalent to holding it wrong
        • Keyframe 19 hours ago
          You both can be right at the same time.
      • bigstrat2003 19 hours ago
        There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.
  • lkramer 21 hours ago
    I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.
    • Cyan488 20 hours ago
      I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.

      I think a PR "in flight" shouldn't have been closed like that.

      All this will do is push out developers like you that honestly disclose, and instead people will now just lie.

    • sedan_baklazhan 20 hours ago
      What is actually the meaning of “handcrafting” here?
      • throwaway74628 20 hours ago
        My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.
        • thereisno 18 hours ago
          "taking responsibility" doesn't actually mean anything, it's just empty words
          • simonra 17 hours ago
            Putting any specific string in an "author(s)" or "committer" field is likewise a token gesture that doesn't affect or alter the content. The impact, just like promises of responsibility, would still solely be in biasing the reader, clouding the evaluation.

            Unless you believe there is substance in social interactions over time, like trust. But then it would be a socially weird move to dismiss promises of responsibility out of hand instead of picking up the invitation to build trust if that's what's perceived to be lacking.

          • throwaway74628 10 hours ago
            We are not the same.
      • lkramer 17 hours ago
        What I meant by it was that I read every line of code generated, made sure I understood its purpose and either manually rewrote it if I felt there was a better way or asked Claude to do it. E.g. there was a bug where the password could end up in a configuration labelled as a username. Claude's initial fix was to simply exclude 'username' in a for loop, but I asked it to find examples in similar code in other codebases to see what the best practice was and ended up basing to fix on what is done in Gnome.
  • sippingabonedry 20 hours ago
    Completely performative.

    You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?

    • driverdan 19 hours ago
      It's not. Read the article, they have a good reason for banning it and it's not because they hate LLMs.
    • _heimdall 20 hours ago
      It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.

      I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.

    • 999900000999 19 hours ago
      I agree.

      PopOS is Ubuntu with extra problems. Ubuntu itself is fine, but then PopOS adds weirdness.

      Cosmic has been in beta for how long ?

      • plqbfbv 18 hours ago
        > Cosmic has been in beta for how long ?

        I daily drove the alphas before the betas, and of course there were a couple rough edges. But I had a minimal working desktop instead of sway or KDE/Gnome (too heavy).

        It has been releasing non-betas for a good year now, and it's been a smooth sailing.

      • wookmaster 14 hours ago
        Huh? Cosmic is out, there's some bugs still but its pretty nice.
      • duped 19 hours ago
        It's been out of alpha/beta for almost a year
    • novafunc 19 hours ago
      I don't think it's performative.

      They aren't saying that AI produces bad code or is terrible for the world in some way.

      It's mainly just resulting in a lot of PRs that they don't have enough time to review or features they don't plan to add.

    • schipperai 18 hours ago
      It sounds like what they are primarly aiming for is introducing backpressure in their own review pipelines.
    • HexDecOctBin 20 hours ago
      This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.
      • brokencode 19 hours ago
        There was plenty of bad code long before LLMs came around. Let’s be real.

        Being hand written is no guarantee of high quality, just like using LLMs is no guarantee of low quality.

        • bigstrat2003 19 hours ago
          Using LLMs is not a guarantee of low quality, but it does make low quality vastly more likely.
          • brokencode 18 hours ago
            Eh, not really. It just lowers the barrier for lazy people who never would have tried contributing before.

            I’ve read a lot of code hand-written by programmers. By smart and hard working people. And I know from experience that the code the average programmer writes is not great either.

      • Fr0styMatt88 19 hours ago
        Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).

        This is just really silly.

        The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.

  • ItsMattyG 21 hours ago
    I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?
    • PorciiVorbesc 21 hours ago
      Probably because pop_os and Cosmic are so niche and their market share so insignificant, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry to make their time and effort worth it. See the Arch AUR attacks, for perspective.

      I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.

      But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.

      The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.

      • satvikpendem 20 hours ago
        You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.
        • PorciiVorbesc 19 hours ago
          >You only need one bad actor.

          If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.

          • satvikpendem 19 hours ago
            Of course. I never said life doesn't go on or it's a worthwhile risk to care about, not sure how you got that from my comment.
            • PorciiVorbesc 19 hours ago
              > not sure how you got that from my comment.

              Because your comment didn't disprove that Cosmic DE isn't too niche for bad actors to get involved.

              • satvikpendem 19 hours ago
                It doesn't have to be niche (or not) to have bad actors. That doesn't make it likely of course but it is possible, and increasingly more so in the age of AI when you can simply point it at multiple projects in parallel.
                • PorciiVorbesc 18 hours ago
                  >It doesn't have to be niche (or not) to have bad actors.

                  Then they shouldn't reject AI aids to help them patch vulns found by bad actors faster, no?

                  • satvikpendem 17 hours ago
                    I never said they should. I think you're reading too much into my comment, the point was something being niche doesn't prevent it from being threatened.
      • plqbfbv 18 hours ago
        > Probably because pop_os and Cosmic are so niche and their market share so insignificant

        Consider that it's packaged for many well-known distros, so pop_os install base alone doesn't tell the whole story: https://system76.com/cosmic/download

      • badc0ffee 20 hours ago
        It does show up on top 5 lists for Linux desktop distros quite a lot, and COSMIC is quite unique, so I suspect a lot of people are at least trying it.

        (I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)

        • satvikpendem 18 hours ago
          Source on this?

          Edit: I see you posted elsewhere.

        • PorciiVorbesc 20 hours ago
          >It does show up on top 5 lists for Linux desktop distros quite a lot

          1) Firstly, your source plase? My research according to Google Gemini 3.8 Pro shows top 5 DEs are as follows:

            +---+---------------+------------+-------------------------------------+
            | # | Environment   | Est. Share | Primary Ecosystem / Defaults        |
            +---+---------------+------------+-------------------------------------+
            | 1 | GNOME         | 45% - 50%  | Ubuntu, Fedora, Debian, RHEL        |
            | 2 | KDE Plasma    | 25% - 30%  | SteamOS, openSUSE, Kubuntu, Manjaro |
            | 3 | Cinnamon      |  8% - 12%  | Linux Mint flagship                 |
            | 4 | Xfce          |  6% - 9%   | MX Linux, Xubuntu, low-spec PCs     |
            | 5 | MATE          |  2% - 4%   | Ubuntu MATE, Mint MATE (GNOME 2)    |
            | - | Others / WMs  |  3% - 5%   | Hyprland, i3, Sway, LXQt, Budgie    |
            +---+---------------+------------+-------------------------------------+ 
          
          So then COSMIC isn't even TOP 5, for this to be a major target by market share as originally claimed.

          2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.

          3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.

          • yugoslavia4ever 20 hours ago
            You're being way too defensive. GP replying to you was clearly trying to have a conversation, not attack your knowledge. Chill.
            • PorciiVorbesc 19 hours ago
              >You're being way too defensive.

              "Way too defensive" how? By asking and bringing data for my PoV?

              >GP replying to you was clearly trying to have a conversation

              As am I, except I ask for, and also bring data to back up my PoV, instead of vague opinions.

              >Chill.

              Where am I not being chill?

              • kyubey 18 hours ago
                To reflect some sentiment from another comment you posted - we're not your unpaid tone auditors. Learn to have normal discussions.

                For funsies, I threw two different DE-by-marketshare inquiries at Gemini 3.8 Pro in separate sessions and it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.

                • PorciiVorbesc 18 hours ago
                  >Learn to have normal discussions.

                  Please show me what part of what I said before was not "a normal discussions" according to you?

                  >it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.

                  And that disproves me how exactly? I originally showed "Cosmic is NOT a TOP5 DE", and the LLM data you posted also shows that IT IS indeed NOT a top 5 DE.

              • cindyllm 18 hours ago
                [dead]
          • satvikpendem 20 hours ago
            Don't post AI generated content, if needed post the actual source.
            • PorciiVorbesc 19 hours ago
              >Don't post AI generated content

              Do you have a better source than GP?

              >if needed post the actual source

              Define "actual source"? In good faith, I mean.

              Where else do you get this information that's, quote, "actual source"?

              • satvikpendem 19 hours ago
                As in the study or survey or analytics segment where the data was collected to produce the report.
                • PorciiVorbesc 19 hours ago
                  "The analytics", is the publicly available training dataset of the LLMs that share the same common opinion on that DE market share.

                  What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.

                  If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.

                  • satvikpendem 19 hours ago
                    Point at the public dataset then if it's so public. You really have no idea how the LLM is parsing the data and could easily be hallucinating. That you find being called out on this as some sort of insult is quite telling and frankly funny, what a strong reaction to someone asking for a basic source when the burden of proof is on you to prove (or at least show a source that) these stats are right, not on anyone else to disprove AI bullshit.
                    • PorciiVorbesc 18 hours ago
                      Sorry, it's not my job to provide for you the things that you demand from me at your whims, because the original comment I replied to with LLM data also did not bring peer-reviewable information as proof that COSMIC is a TOP 5 DE, and yet you did not demand proof from them in that case. Why is that?

                      Did you just blindly trust the opinions of others on this topic, and yet I'm the one who has to provide peer-reviewable data for you to back-up mine? Sorry, I'm not your unpaid lackey. Try to formulate a better (counter)argument for why their baseless opinion is right, but my LLM backed up opinion is wrong, if you wish for a even-footed good-faith argument.

                      • satvikpendem 18 hours ago
                        I ask them for the same source (and looks like they provided it), your comment wasn't special. LLMs however are especially less trustworthy, that's why. It's not my job to educate you on why they are, as you say, and why people aren't trusting your comments.
                        • PorciiVorbesc 17 hours ago
                          >I ask them for the same source (and looks like they provided it)

                          I also saw it now. That blog is not a representative ground truth, but just another opinion piece, which I can respect as an opinion of the blog's user base, but I can't take as an accurate real world statistic, same how aggregate opinions you read on HN are not representative of the actual real world.

                          I hope you can understand my PoV. You can also disagree if you want, but you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.

                          • satvikpendem 17 hours ago
                            In theory if you brought a better source than that person then I'd agree with you, but,

                            > you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.

                            actually I can say it's wrong or likely to be wrong especially if it doesn't cite the sources it uses. And if it does, then just paste the sources here instead of the LLM output. It is also unknown where it got the info and as someone else said, you ask it two different times and it gave two different answers, thus it is unreliable.

          • badc0ffee 19 hours ago
            Sorry for being unclear. I meant it shows up as a top 5 distro recommended by reviewers. This kind of thing: https://linuxblog.io/best-linux-distro/
            • PorciiVorbesc 19 hours ago
              OK, but what's the sample size of that and who's measuring it? I never heard of that blog or took part in that poll. So how is that blog link the yardstick but mine is not?
              • nvme0n1p1 18 hours ago
                Because that blog actually talked to real people and did real work, and yours is just a hallucinated list from one of the me-too LLM vendors
                • PorciiVorbesc 18 hours ago
                  >Because that blog actually talked to real people and did real wor

                  How did you verify that those people from the blog are "real"?

                  I also talked to real people for my own data, case in point, I asked my mom and dad which linux DE is most used, and the results came out different. Which "real people" are the ones representative for the ground truth of Linux DE sahre?

                  > and yours is just a hallucinated list from one of the me-too LLM vendors

                  How do you know it's hallucinated? Ask the LLM the population of your country? Is the answer mostly accurate or is it hallucinated in an inaccurate way?

                  Aren't LLMs just outputting the highest statistical probability from the aggregate of their scraped data, which in this case would be including opinions on Reddit, and every website and blog on the entire internet (including that random one posted by badc0ffee) on the Linux DE uusage topic, making it a more accurate real-world representation than just a single random blog?

                  You can call it "hallucinated" if you want, but that doesn't mean it's not accurate. I asked for proof that my answer was inaccurate, not that it was "hallucinated", those are two different things, and your argument didn't prove it was inaccurate nor did it prove it was hallucinated. Would you like to try again?

      • iugtmkbdfil834 20 hours ago
        << Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.

        That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.

        << I think even amongst the HN and Linux userbase, pop_os is still niche

        I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?

      • DelightOne 20 hours ago
        It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.
        • PorciiVorbesc 20 hours ago
          Sorry, I don't understand what you mean by this, can you elaborate pls?
          • DelightOne 19 hours ago
            Its easier for an LLM to find vulnerabilities in a project with less usage, and those vulnerabilities will stay open longer, making them much cheaper to attack to keep the door open.
            • PorciiVorbesc 19 hours ago
              What's the point of attacking projects that almost nobody uses?

              Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?

              Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.

              That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.

              • DelightOne 19 hours ago
                If it costs you as good as nothing, you might as well do it.

                Previously, time was the most precious resource. Now its tokens, and more cheaply at at.

                • PorciiVorbesc 19 hours ago
                  Offensive security employees, tokens, and peoples' time are still a finite resource that get allocated based on target priorities and operational end-goals, even by state actors.

                  If you assume Mossad and NSA are Token-maxxing every single niche FOSS project out there to cast as large as possible fishnet on hacking all Average Joes on the planet just in case, then maybe using Mozilla and MacOS gets you hacked too, maybe even visiting HN and commenting here gets you hacked by some zero days you don't yet know.

                  Where does this open-ended paranoia argument end?

    • fwlr 19 hours ago
      So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.
    • layer8 20 hours ago
      Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.
    • altcognito 20 hours ago
      Irony is that they will get the benefits from upstream projects (like the Linux kernel) that does accept AI inputs.
    • VCFundedGenYer 20 hours ago
      AI finds a lot of "vulnerabilities" but most are fake, untested, or not actually vulnerabilities.

      Reminder that AI is quite stupid.

      • novafunc 19 hours ago
        It may have a high false positive rate, but at the speed AIs can review code, there's still plenty, of real vulnerabilites mixed in with the garbage.

        I'll trust the words of groups like curl (https://daniel.haxx.se/blog/2026/06/10/a-human-in-control/), Linux, and even the infamously anti-AI Gnome (https://blogs.gnome.org/mcatanzaro/2026/10/02/the-era-of-sof...) that AI is finding real vulnerabilities and you're your project a disservice by ignoring them.

        Edit: Though Greg did recently have a talk (that I skimmed) where he was a little reserved on LLMs: https://www.youtube.com/watch?v=NnV_cWeoo5Q

      • dumberquestions 20 hours ago
        I think this is just plain denial, AI has found many high severity vulnerabilities.
        • zdragnar 20 hours ago
          There's plenty more false positives than actual finds. There are still actual finds, but that doesn't change all the false positives.
          • DaSHacka 19 hours ago
            This is why you run a second agent to verify any assertions.
            • zdragnar 13 hours ago
              It'd be great if people did that. They don't. Open source projects with limited budgets shouldn't have to spend money from those limited budgets (or personal funds!) filtering through what amounts to spam. That's why there was a slew of open source projects that turned off issues on github and stopped taking outside contributions- they literally couldn't afford to keep up with the nonsense lazy bums were sending their way.
    • vorticalbox 20 hours ago
      The issue is “ai generated code” using ai to find bugs/0 day and manually writing a fix would (I assume) be allowed under the new rules.
    • barbazoo 21 hours ago
      I'm assuming they still use AI to find vulnerabilities, just not fix them?
  • YuechenLi 18 hours ago
    Ok, this may be controversial, but LLM code tokens aren't free, and I run out of my weekly allowance pretty regularly just from doing some fairly heavy projects, so I don't understand why somebody would ever want to spend their own money to make bad PRs on purpose, and I like to assume good intentions from people unless proven otherwise, which means a near blanket ban for LLM authored code for these big open source projects just seemed a bit extreme to me, when the core issue seemed to be that review process/policy should change with the times.

    For example, I was helping work on an open-source game engine earlier this year with a longstanding text rendering bug dating back to around 2021 that prevents the engine from being production ready, which the community and myself have developed extensive workaround for. So, one day I've finally said enough and got Claude to debug it. It took Claude 10 minutes to find the bug, it was 3 lines of code change in the renderer (yes, three).

    So, I wrote up the regression tests, documented the bug and opened up a PR for the fix, thinking it'll get merged in like less than a week and then we can all move on. The maintainers received it fairly well on the PR, but the PR sat there for nearly 6 months, unmerged, until it finally closed from a bad squash upstream. I'm pretty sure the bug is still there too.

    And as a side note, I would be ecstatic if someone wants to contribute to my Github projects with their AI.

    • northstar702 15 hours ago
      Feels like AI on autopilot running OSS projects is the dream everyone is trying to get to. Tokens will never be free. Anthropic needs to go IPO. You make an interesting point about personal spend, but bad PRs are unlikely to be on purpose. But the discipline to review/clean up the slop might not be there?
      • YuechenLi 14 hours ago
        True. I personally use LLMs much more ambitiously to help with my own coding projects, but I've always been very disciplined about planning ahead and task breakdown so that the LLMs can get it right the first time so there won't be any need to fix the code afterwards.

        Other people's open-sourced project is different though, so I tend to check/test the code myself much more rigorously when contributing to other people's repos than if it is my own projects.

  • winrid 20 hours ago
    I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol
    • fullstackwife 18 hours ago
      Issue is entrenchment in the old SWE world that ceased to exist somewhere in August this year, and using AI to mechanize the traditional workflows. Also people with "I need to eyeball each char in PR diff" attitude, which is fucking unproductive at this point.
  • i_love_retros 18 minutes ago
    Good! Anything that isn't a not very important crud web app should not allow AI generated code anywhere near it!
  • cagey 13 hours ago
    I had a PR closed because of this, one that fixes something that, had it not been fixed (by Claude, based on me noticing and clarifying the failure mode and verifying the fix), would have led me to abandon at least COSMIC, if not the distro. But this has me wondering how COSMIC, which had reached 1.x before I started using it, can mature, and whether I'll remain. I started on Pop!_OS with its COSMIC-less 22.04 version; I chose it as a fairly popular distro having Ubuntu stability while offering a leading-edge kernel. Point being, COSMIC was never the draw for me, and I held off upgrading to 24.04 for a LONG time out of concerns about bugs & stability; eventually 22.04 started experiencing ibus-related problems that forced my hand. Since I am (a) not knowledgeable in Rust, nor (b) conversant in the protocol at the core of this bug, I would never have had the gumption or time to understand and fix the bug. And since the PR is closed, it seems unlikely to even serve as a bug report.

    In the meantime, others have experienced this bug and are waiting for the maintainers to notice and fix it. I'll just continue to run my fork (so far it has been no burden at all) while I (slowly) decide what direction to take.

    • mkj 3 hours ago
      Maybe you could report it as a normal bug report, by hand?
  • teekert 19 hours ago
    "... many of the AI contributions were not planned and showed little understanding of the software architecture. So the team wants to "prioritize working on contributions from our own team and regular contributors.""

    Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.

  • northstar702 16 hours ago
    there is an ongoing thread here on a similar topic from an AWS expert (former AWS CTO)

    "Apparently, AI doesn't lead to positive results in all software development teams. Customers are asking me whether they should slow down the adoption of AI in their teams.

    How do you respond to that? Yes? No? "

    https://www.linkedin.com/feed/update/urn:li:activity:7510679...

    Feels like part of it is a learning habit, getting proficient in use of the tools (AI agents) themselves, and better workflow around it, but it remains AI is not perfect yet?

  • yegle 19 hours ago
    This would just push people to maintain their own fork. If I already have an agent to investigate and fix a bug and able to send a PR, the added cost of maintaining a local fork is minimum.

    In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.

  • VCFundedGenYer 20 hours ago
    Thank god. Tired of these orgs blindly and stupidly allowing them (Debian, cough).
  • sdcfgy 20 hours ago
    Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.

    Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.

    • driverdan 19 hours ago
      A 100kloc PR is unacceptable regardless of the source. That needs to be broken down into reviewable chunks.
  • onesandofgrain 19 hours ago
    As they should. If you can tell it's AI, it shouldnt be part of any codebase.
  • snvzz 7 hours ago
    Very quickly they'll find themselves with most packages frozen to old versions, including the kernel itself.

    Shipping them full of bugs which LLMs already fixed upstream.

    Fortunately, as it is an open source project, if anyone actually cares about the distribution, they'll fork it and continue advancing it using modern tools.

  • holoduke 20 hours ago
    Doing that will put you put of the market. I am certain of that.
    • lumpysnake 20 hours ago
      There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.
      • baq 20 hours ago
        it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.
        • bakugo 19 hours ago
          This is straight up cult behavior. "You WILL assimilate or we WILL kill your project"

          PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.

          • baq 19 hours ago
            it isn't even me not liking it (I don't think about it at all frankly, I have nothing to like or dislike) - I think it's irresponsible policy. I won't be using something that doesn't accept automated security patches from defender AIs.
            • bakugo 18 hours ago
              Okay, so don't use it. Again: it's an open source project, its maintainers don't owe you anything, nor is anyone claiming you have to use it.
              • baq 17 hours ago
                I won't and I recommend anyone currently using it to reconsider due to the security posture the project necessarily and unfortunately has as a direct consequence of its policies.
  • zzzeek 21 hours ago
    banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
    • nicoburns 20 hours ago
      Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.

      My repos probably don't see as much traffic as SQLAlchemy though.

    • nonethewiser 21 hours ago
      > If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?

      Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.

    • baq 20 hours ago
      rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)
      • grumbel 18 hours ago
        I'd still reject them. It's the bug reports and feature requests that are valuable. When you have an AI yourself, there is little point in having somebody else let their AI implement them, that just creates a lot of risks and unknowns for no benefit.
      • zzzeek 16 hours ago
        yeah a lot of LLM PRs are pretty good, but still need changes, and still didnt come from my own prompting which would have got them more exactly where I want them, so it's again, I have to put messages on a PR and wait for someone somewhere to see them and act on them. That friction is a huge waste of time if they're just prompting their own robot. I have the same robot right here and I usually use opus 5.x which is usually better than what they're using.
        • baq 3 hours ago
          I think it’s perfectly ok to just tell your robot how to get a good pr to a state you like and merge it then. You should of course let folks know that this is the policy, but ultimately if it solves their problem the way you like it and you didn’t have to spend your tokens on it everybody wins…?
          • zzzeek 2 hours ago
            i use claude max and ive yet to get even 30% into my quota for it. so as long as that kind of deal lasts i dont think much about tokens.

            if i did have to think about tokens I use something like together.ai with an open weight model like glm 5.3 (which ill sometimes use to review a claude change for something intricate). glm 5.x is just very chatty though

    • api 20 hours ago
      Maybe the policy should be: no AI PRs except from established contributors who have been vetted?

      So basically you're saying you reject drive-by PRs.

      • whateveracct 20 hours ago
        drive-by PRs that are handwritten are presumably okay
  • mahboi 20 hours ago
    I still wanna see what Slop OS looks like. Go full AI spam making a Linux distro from the kernel up.
    • slig 19 hours ago
      It might get millions in founding.
  • aronaxe 12 hours ago
    [flagged]
  • iluvcommunism 20 hours ago
    [dead]
  • CurbStomper4 20 hours ago
    [dead]
  • whatsThisBtn4 20 hours ago
    Well this project is dead then.

    I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.

    Unironically.

    Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.