GitCode, a git-hosting website operated Chongqing Open-Source Co-Creation Technology Co Ltd and with technical support from CSDN and Huawei Cloud.

It is being reported that many users’ repository are being cloned and re-hosted on GitCode without explicit authorization.

There is also a thread on Ycombinator (archived link)

  • dan@upvote.au
    link
    fedilink
    English
    arrow-up
    95
    arrow-down
    14
    ·
    edit-2
    6 months ago

    I don’t understand why this is a bad thing? Open source code is designed to be shared/distributed, and an open-source license can’t place any limits on who can use or share the code. Git was designed as a distributed, decentralized model partly for this reason (even though people ended up centralizing it on Github anyways)

    They might end up using the code in a way that violates its license, but simply cloning it isn’t a problem.

    • BlueMagma@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      36
      arrow-down
      4
      ·
      6 months ago

      I expect it’s going likely to be used to train some Chinese AI model. The race to AGI is in progress. IMO: “ideas” (code included) should be freely usable by anyone, including the people I might disagree with. But I understand the fear it induces to think that an authoritarian government will get access to AGI before a democratic one. That said I’m not entirely convinced the US is a democratic government…

      PS: I’m french, and my gov is soon to be controlled by fascist pigs if it’s not already, so I’m not judging…

      • dan@upvote.au
        link
        fedilink
        English
        arrow-up
        12
        arrow-down
        3
        ·
        6 months ago

        I expect it’s going likely to be used to train some Chinese AI model.

        Even if they do that, the license for open source software doesn’t disallow it from being done.

        • sugar_in_your_tea@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          13
          ·
          6 months ago

          It certainly can. Most licences require derivative works to be under the same or similar licence, and an AI based on FOSS would likely not respect those terms. It’s the same issue as AI training on music, images, and text, it’s a likely violation of copyright and thus a violation of open source licensing terms.

          Training on it is probably fine, but generating code from the model is likely a whole host of licence violations.

          • dan@upvote.au
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            4
            ·
            6 months ago

            Most licences require derivative works to be under the same or similar licence

            Some, but probably not most. This is mostly an issue with “viral” licenses like GPL, which restrict the license of derivative works. Permissive licenses like the MIT license are very common and don’t restrict this.

            MIT does say that “all copies or substantial portions of the Software” need to come with the license attached, but code generated by an AI is arguably not a “substantial portion” of the software.

            • sugar_in_your_tea@sh.itjust.works
              link
              fedilink
              English
              arrow-up
              8
              ·
              6 months ago

              code generated by an AI is arguably not a “substantial portion” of the software

              How do you verify that though?

              And does the model need to include all of the licenses? Surely the “all copies or substantial portions” would apply to LLMs, since they literally include the source in the model as a derivative work. That’s fine if it’s for personal use (fair use laws apply), but if you’re going to distribute it (e.g. as a centralized LLM), then you need to be very careful about how licenses are used, applied, and distributed.

              So I absolutely do believe that building a broadly used model is a violation of copyright, and that’s true whether it’s under an open source license or not.

    • barryamelton@lemmy.ml
      link
      fedilink
      English
      arrow-up
      24
      arrow-down
      1
      ·
      6 months ago

      The code needs to maintain the copyrights and authors. They are “mirroring” usernames into their own domain, with mails that dont correspond to the original authors, stealing their contributions.

      • dan@upvote.au
        link
        fedilink
        English
        arrow-up
        5
        ·
        6 months ago

        with mails that dont correspond to the original authors,

        Oh! I didn’t realise this. Do you have an example?

      • Aceticon@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        arrow-down
        1
        ·
        edit-2
        6 months ago

        That would make it plagiarism, which ethically is a whole different matter than merelly copying that which is free to copy.

    • Kayn@dormi.zone
      link
      fedilink
      English
      arrow-up
      19
      arrow-down
      2
      ·
      6 months ago

      I’m seeing this misconception in a lot of places.

      Just because something is on GitHub, doesn’t mean it’s open source. It doesn’t automatically grant permission to share either.

        • Kayn@dormi.zone
          link
          fedilink
          English
          arrow-up
          3
          ·
          6 months ago

          Correct, you are allowed to click the “fork” button and nothing else. You’re still not allowed to download, use, modify, compile or redistribute the code in any way that doesn’t involve the “fork” button.

      • Grimm665@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        5
        ·
        6 months ago

        It may not be de jure open source, but if the code is posted publicly on the internet in a way that anyone can download and modify it, it sort of becomes de facto open source (or “source available” if you prefer).

        • JackbyDev@programming.dev
          link
          fedilink
          English
          arrow-up
          5
          ·
          6 months ago

          Please don’t muddy the water with terms like this. Something is open source if and only if it has an open source license.

    • ZILtoid1991@lemmy.world
      link
      fedilink
      English
      arrow-up
      16
      arrow-down
      2
      ·
      6 months ago

      I personally don’t care if someone “steals” my code (Here’s my profile if you want to do so: https://github.com/ZILtoid1991 ), however it can mean some mixture of two things:

      1. China is getting ready for war, which will mean the US will try its best to block technology, including open source projects.
      2. China is planning to block GitHub due to it being able to host information the Chinese government might not like.

      Of course it could mean totally unrelated stuff too (e.g. just your typical anti-China and/or anti-communist paranoia sells political points).

      • dan@upvote.au
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        1
        ·
        edit-2
        6 months ago

        US will try its best to block technology, including open source projects.

        You can’t block open source projects from anyone. That’s the entire point of open source. For a license to be considered open-source, it must not have any limitations as to who can use it.

        • irreticent@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          6 months ago

          You can’t block open source projects from anyone.

          I think they were referring to blocking GitHub from public access. In the event of a world war I could easily see Microsoft obeying the order to shut down GitHub.