hi Spencer, I am still not sure if things are well understood in this field, let me try to explain what is happening during that training, pls; that machines neural net (think about it as some kind of "brain simplified model", in this case based/focused on LLM concept, so understanding the meaningof languages, spoken or formal, programming, far better defined...) is just really "learning" from others open-sourced work, really like anybody else on the planet, while reading the source, for learning purposes - which happens, always, but in human limited brain its not so effective/deep, sure, somebody digs deeper and focuses on some repo and learns a lot from this one case probably a may even fork it and extend or not, and/or do some new work inspired by code from he/she somehow learned something, like from books or courses, anything around and anybody else is doing... fact is these machines are bigger than human brain and really good in this LLM field these days ... in case there is scanend 175 000 000 repos, you can calculate what theoretical percentage of someones contribution to this sea of knowledge is ... but it is NOT in any way "stealing actual code" ... the neural net is in fact observing/comparing/evaluating all existing work, all different forms some good some worse, some bad, with all the context of impact and usage and likes and stars and forks putin ALL this into "weights" on the neural network (those bilions of parameters, say aka neurons and their inputs ... imagine neuron as some kind of "analog computer fuzzy-logical gate/opamp" ... ya, this is what we all have in brains, usually, and the LLM models are really learning the same way as human brain ... ya, its faster, its fast as hell, in fact, it may be good or bad, depends on intents/usage, depends on who uses it then for what ... sad thing about open-source is that NO licenses are preventing bad usage of code, I asked the machine for these things too, but no, it looks mostly like open-source principle is intentionally abused to be "scrapped" for possible use in any way, even for weapons (generall consensus they are not abd thing, especially today, but weights of this were changing in peace times) development (which is something I would personally try to prohibit, but its legally almost impossible ... all the licenses are designed the way, that this prohibition is impossible, almost ... really, "bad usage" (and it all is also "relative", based on subjective point of view) is hard to prevent, HERE I dont like it ... the neural netoworks (LLMs) are not the culprit of bad things, they can do lot of good thing (mostly TEACHING people, if they are prepared to LEARN), as aggregated source of knowledge available to ENTIRE humanity ... take it as your source codes are only tiny tiny drop of water in entire sea, again 175M repos, vs few of ours ... each one is adding only tiny part of his own approach to this knowledge (where I opt-in, intentionally, expecting drastically tiny weights though, sure) ... the machine is not grabbing entire sources and using even "single exact line" from any those scanend repos ... the knowledge is "averaged" inside the built neural network, the exact source is immediatelly lost in the huge ocean but depending on quality of work it somehow adds or subtracts something from some parameters (neurons inputs) and their weights to help that neural netowrk to "emerge" OWN solution /approach for some probelm, based on learned things (and LOTs of them.... sure, far more than any single human, sure) ... it may sound also scary, sure (and it REALLY is, no doubt, in case of usage of such knowledge can be targeted by some "bad human actor" to bad direction, used for bad things (but at the machine level are ALSO intentional guardrails today, able to detect any such bad behavior...) ... but if you believe in humanity, that, thumans are generally good, and that we as civilisation hopefully learned something during history, and we have some experience how bad things start and where/how they can end (me noting mostly WWII result of nazi idiocy rise, but in fact all wars back in history, sure) ... then when you want to believe in good people and in humanity as a whole, then you can even HELP by your contribution and good intents, but only by that tiny drop of water in entire ocean ... the machine then doesnt know or reference your work (ya, now it may "remember" and link to it if its some weighted interesting solution described by constrains of problem solved, when the "machine brain" decides its good to refer to it in links) ...and as it is coding, it is formal, it is testable ... now using "agentic loops", the machine can test the code in real environment, see results, read error reports (just like a human) and react accordingly, try to fix problem (it doesnt do blind tries, its good at it, it saves time for humans, it just now already "really thinks" about problems, and by adding coded processes and "standard operating procedures" in SW developemnt, it is at quite high level already, ... in fact far better that even some "blind juniors" .... and this power is available to anybody to help to do his actual new work ... its not generally bad and imho it is not "stealing" ... the principle of open-source was to be things available to others, at least visible, ... then it depends on licenses what all can be done with it ... imho, intentionally no licenses are preventing making weapons, ... ya even LLMs may be one of them, too ... may be, but it is not definitive, it may be also good thing ... it all depends, what is "meaning of life";
(this was obvously not generated by LLM, I swear)
Petr (apws repos, drop in ocean)