GitHub code scraping

129 views
Skip to first unread message

Spencer

unread,
Jul 25, 2026, 8:31:30 AMJul 25
to RC2014-Z80 GoogleGroup
Some of you may be aware of an AI company called Huggin Face.  They are currently scraping up as much data as they can to train their model.  This includes ALL of GitHub.  

To try and keep this message on topic for the RC2014 Google Group, I am just concertning myself with the RC2014Z80 or other RC2014 related repos*, and trying not to stray in to the morality of AI scraping like this. This does have an Opt-Out facility (Don't get me started on how opt-out rather than opt-in is any form of concent!), so I have requested the removal of all RC2014Z80 data from their servers.

If you have any GitHub repos with RC2014 projects and you don't feel like just giving these over to them, the opt-out page is here and is very quick and easy to do

Spencer

* This does, of course, apply to EVERY repo

Doug Jackson

unread,
Jul 25, 2026, 9:53:50 AMJul 25
to rc201...@googlegroups.com
It is a complex topic.

Open access is just that.  Open.

I moved to GitLab when Micro$oft brought GitHub.  On the simple basis that if you dont pay for it, you are the product....

But there is also the reality that your product may just as likely train a class of Comp101 graduates as an AI, and we will never know.

I work on the basis that if it's on a public server, it's public.  

Just my 3 cents worth.   

Doug 


--
You received this message because you are subscribed to the Google Groups "RC2014-Z80" group.
To unsubscribe from this group and stop receiving emails from it, send an email to rc2014-z80+...@googlegroups.com.
To view this discussion, visit https://groups.google.com/d/msgid/rc2014-z80/OV1C2Q-UgFQ3MtxKQXVQBFdfUIuOW0DDgIq3hZL9mEk3WTgP9xrfWAvK3rCl-si1zJjxm_t1xL2znDsHASM7PzqwZoUe4KCERbE7gOThwP8%3D%40sowen.com.

Mark T

unread,
Jul 25, 2026, 12:28:26 PMJul 25
to RC2014-Z80
Do you trust them to take notice of the opt out and not use any information provided as additional input?

Spencer

unread,
Jul 25, 2026, 12:55:09 PMJul 25
to rc201...@googlegroups.com
Frankly I don't trust them one bit. But short of going to their data center and standing over a techie whilst watching him delete it from the current live data and all backups, I'm not sure what other options there are.

Also, I suspect that this opt-out is only for this particular data grab. I'm sure there'll be more.

Spencer

--
You received this message because you are subscribed to the Google Groups "RC2014-Z80" group.
To unsubscribe from this group and stop receiving emails from it, send an email to rc2014-z80+...@googlegroups.com.

Steven P

unread,
Jul 25, 2026, 2:58:41 PMJul 25
to RC2014-Z80
Unfortunately, the opt-out (assuming the take note of it) would only have an impact if no-one else has made a folk of your source code. I'm sure the AI is right now reviewing that opt-out file and forking... :-)

Gary Hammond

unread,
Jul 25, 2026, 4:55:08 PMJul 25
to RC2014-Z80
For my code that I have shared publicly, I am not that fussed. If I have open sourced it, then it's free for anyone to use, private or commercial. My logic is that I have gained so much from other peoples willingness to open source, that I feel obligated to do the same in return.

For code that you care about and don't want commercial interests in using, then you need to find a different way of sharing your code. Not sure how that is going to work, because even when commercial interests finds a way around that and use it anyway, there is no repercussions for them. None of us likely have the financial wherewithal to sue them over it even if we do find out they have used it. 

Maybe it's back to old fashioned mailing lists and FTP servers. (I am all for that BTW). The problem is how to get the word out without tipping off the big boys.

John W. O'Brien III

unread,
Jul 25, 2026, 6:23:22 PMJul 25
to rc201...@googlegroups.com
different variations of Conway's Life (hex output rolling up 4 cells
into one digit and then coloring that digit in the terminal) a paging
text reader written in DX-Forth, a bare bones procedural generated
dungeon crawler, a few versions of Tetris...
all written for either the SC131 (and other ROMWBW machines, or the
Pocket386 (of which one of mine has bad memory i think, my "Prime386"
stress test program wouldn't run... lol... found the possible hardware
fault... same code, same TP7 compiler work flawlessly in Dosbox)

M$ can have CAT.fth, call it compensation for me sailing the 7-seas in
college.... and still...

https://github.com/inerlogic/retro-code

dean.ne...@gmail.com

unread,
Jul 25, 2026, 8:26:32 PMJul 25
to RC2014-Z80
Thanks for the link Spencer - just submitted my opt-outs.

I guess we should be grateful that at least we have the opt-out for hugging-face - I am sure MS has used all github projects for training their copilot --- I remember the controversy of early versions of copilot seemingly reproducing peoples code verbatim- including weird comments - without any regards to licencing/copyrights.

A reason I started to moved to codeberg.org.  Their members have recently voted to confirm new policy regarding use of AI - disallowing scraping for training and also disallowing projects that are mostly vibe coded. (https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html)  

Although the main reason was to just move off a mega corporate owned platforms - and use a community owned/managed platform - more like what i had imagined the internet was to become.

Dean.

7alken

unread,
Jul 26, 2026, 9:30:14 AM (14 days ago) Jul 26
to RC2014-Z80
hi Spencer, I am still not sure if things are well understood in this field, let me try to explain what is happening during that training, pls; that machines neural net (think about it as some kind of "brain simplified model", in this case based/focused on LLM concept, so understanding the meaningof languages, spoken or formal, programming, far better defined...) is just really "learning" from others open-sourced work, really like anybody else on the planet, while reading the source, for learning purposes - which happens, always, but in human limited brain its not so effective/deep, sure, somebody digs deeper and focuses on some repo and learns a lot from this one case probably a may even fork it and extend or not, and/or do some new work inspired by code from he/she somehow learned something, like from books or courses, anything around and anybody else is doing... fact is these machines are bigger than human brain and really good in this LLM field these days ... in case there is scanend 175 000 000 repos, you can calculate what theoretical percentage of someones contribution to this sea of knowledge is ... but it is NOT in any way "stealing actual code" ... the neural net is in fact observing/comparing/evaluating all existing work, all different forms some good some worse, some bad, with all the context of impact and usage and likes and stars and forks putin ALL this into "weights" on the neural network (those bilions of parameters, say aka neurons and their inputs ... imagine neuron as some kind of "analog computer fuzzy-logical gate/opamp" ... ya, this is what we all have in brains, usually, and the LLM models are really learning the same way as human brain ... ya, its faster, its fast as hell, in fact, it may be good or bad, depends on intents/usage, depends on who uses it then for what ... sad thing about open-source is that NO licenses are preventing bad usage of code, I asked the machine for these things too,  but no, it looks mostly like open-source principle is intentionally abused to be "scrapped"  for possible use in any way, even for weapons (generall consensus they are not abd thing, especially today, but weights of this were changing in peace times) development (which is something I would personally try to prohibit, but its legally almost impossible ... all the licenses are designed the way, that this prohibition is impossible, almost ... really, "bad usage" (and it all is also "relative", based on subjective point of view) is hard to prevent, HERE I dont like it ... the neural netoworks (LLMs) are not the culprit of bad things, they can do lot of good thing (mostly TEACHING people, if they are prepared to LEARN), as aggregated source of knowledge available to ENTIRE humanity ... take it as your source codes are only tiny tiny drop of water in entire sea, again 175M repos, vs few of ours ... each one is adding only tiny part of his own approach to this knowledge (where I opt-in, intentionally, expecting drastically tiny weights though, sure) ... the machine is not grabbing entire sources and using even "single exact line" from any those scanend repos ... the knowledge is "averaged" inside the built neural network, the exact source is immediatelly lost in the huge ocean but depending on quality of work it somehow adds or subtracts something from some parameters (neurons inputs) and their weights to help that neural netowrk to "emerge" OWN solution /approach for some probelm, based on learned things (and LOTs of them.... sure, far more than any single human, sure) ... it may sound also scary, sure (and it REALLY is, no doubt, in case of usage of such knowledge can be targeted by some "bad human actor" to bad direction, used for bad things (but at the machine level are ALSO intentional guardrails today, able to detect any such bad behavior...) ... but if you believe in humanity, that, thumans are generally good, and that we as civilisation hopefully learned something during history, and we have some experience how bad things start and where/how they can end (me noting mostly WWII result of nazi idiocy rise, but in fact all wars back in history, sure) ... then when you want to believe in good people and in humanity as a whole, then you can even HELP by your contribution and good intents, but only by that tiny drop of water in entire ocean ... the machine then doesnt know or reference your work (ya, now it may "remember" and link to it if its some weighted interesting solution described by constrains of problem solved, when the "machine brain" decides its good to refer to it in links) ...and as it is coding, it is formal, it is testable ... now using "agentic loops", the machine can test the code in real environment, see results, read error reports (just like a human) and react accordingly, try to fix problem (it doesnt do blind tries, its good at it, it saves time for humans, it just now already "really thinks" about problems, and by adding coded processes and "standard operating procedures" in SW developemnt, it is at quite high level already, ... in fact far better that even some "blind juniors" .... and this power is available to anybody to help to do his actual new work ... its not generally bad and imho it is not "stealing" ... the principle of open-source was to be things available to others, at least visible, ... then it depends on licenses what all can be done with it ... imho, intentionally no licenses are preventing making weapons, ... ya even LLMs may be one of them, too ... may be, but it is not definitive, it may be also good thing ... it all depends, what is "meaning of life";

(this was obvously not generated by LLM, I swear)
Petr (apws repos, drop in ocean)

Jon Jones

unread,
Jul 26, 2026, 9:56:02 AM (14 days ago) Jul 26
to rc201...@googlegroups.com
Screenshot 2026-07-26 at 14.55.20.png
Reply all
Reply to author
Forward
0 new messages