New License request
Hey!
I was wondering if it would at all be possible to license the Open_SLM_Leaderboard under a new license such as the GPL/AGPL or the Apache 2.0.
While I understand the reasoning for the new license, and I have begun to soften on your position, I still think it would be ideal for the Open_SLM_Leaderboard and freedom as a whole.
I don't really care about what the license is, as long as it follows the Four Essential Freedoms, laid out by Richard Stallman.
I understand that you may not want another fork, which is perfectly reasonable. But, I feel that the SLM community has learned its lesson following the Multivex incident. I'm really starting to view both sides.
I understand if this request cannot be fulfilled, but I figured it would be worth a shot regardless.
Forwarded this to DatDanBoi!
The Multivex Incident™
Forwarded this to DatDanBoi!
Thank you!
To me, the current license feels less like a rejection of open source principles and more like a response to disappointment, idea/work appropriation, and conflict within the community.
To me, the current license feels less like a rejection of open source principles and more like a response to disappointment, idea/work appropriation, and conflict within the community.
it's better to reject open source than get your thing stolen
If I may pitch in my 2 cents, I fail to see the issue. This leaderboard isn't "software" per se (barring semantics), and I don't see any problem with it having its own license. None of the models are bound to that license and can be used freely. I also like the idea that people can't just copy paste results from here (no issues if they ran the evaluations themselves), which can result in staleness and mis-attributions.
Also as not a corporation I don't think this group will be spamming out C&Ds towards those who violate the license in letter but are in compliance with the intended spirit of having it in the first place.
Lets be real here - most of us are making toy models which are practically useless. It is great as a learning experience and see how far we can push the boundaries with the hardware/money we have to spare. Even if there is something interesting, it may be overshadowed by poor dataset selection or training regimen. Don't see the need for a lot of this childishness I have been observing in some of the previous threads regarding this topic (to say nothing of LLM-as-a-User).
To me, the current license feels less like a rejection of open source principles and more like a response to disappointment, idea/work appropriation, and conflict within the community.
I feel that way too. I just still feel that from a moral standpoint, a libre license is better. I’m not saying that Dan is doing anything wrong, I just think it would be better to be under a libre license.
If I may pitch in my 2 cents, I fail to see the issue. This leaderboard isn't "software" per se (barring semantics), and I don't see any problem with it having its own license. None of the models are bound to that license and can be used freely. I also like the idea that people can't just copy paste results from here (no issues if they ran the evaluations themselves), which can result in staleness and mis-attributions.
Also as not a corporation I don't think this group will be spamming out C&Ds towards those who violate the license in letter but are in compliance with the intended spirit of having it in the first place.
Lets be real here - most of us are making toy models which are practically useless. It is great as a learning experience and see how far we can push the boundaries with the hardware/money we have to spare. Even if there is something interesting, it may be overshadowed by poor dataset selection or training regimen. Don't see the need for a lot of this childishness I have been observing in some of the previous threads regarding this topic (to say nothing of LLM-as-a-User).
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
I'm so sorry, but these definitely are toy models (I wonder what po$$ibly the one$ that are not could have in common) and a lot of this "real research" would be shredded apart by any reputable peer review process.
It's great to have pride in one's work and push boundaries to their limits. Overselling things is how spaces like this get toxic.
If I may pitch in my 2 cents, I fail to see the issue. This leaderboard isn't "software" per se (barring semantics), and I don't see any problem with it having its own license. None of the models are bound to that license and can be used freely. I also like the idea that people can't just copy paste results from here (no issues if they ran the evaluations themselves), which can result in staleness and mis-attributions.
Also as not a corporation I don't think this group will be spamming out C&Ds towards those who violate the license in letter but are in compliance with the intended spirit of having it in the first place.
Lets be real here - most of us are making toy models which are practically useless. It is great as a learning experience and see how far we can push the boundaries with the hardware/money we have to spare. Even if there is something interesting, it may be overshadowed by poor dataset selection or training regimen. Don't see the need for a lot of this childishness I have been observing in some of the previous threads regarding this topic (to say nothing of LLM-as-a-User).
Yep, no one is morally required to follow Stallman's definitions by any means, those definitions are better applied to foundational breakthroughs instead of intellectual software work.
Sure, most of these models might look like toys today, but not so much the best ones, and the further you push the line at the scale of the best ones... not to give ideas but those models could be finetuned for good things like targetted research on specific cancer scenarios or really bad automated things that don't require any company's metered safety thresholds.
The smaller you get to do the same things than with the big ones, the more power you have.
OpenAI understood this after releasing GPT-2 and is the reason they no longer open source their models anymore. They would have killed themselves in the process by the way. On the other hand, open source researchers can make incredible discoveries because there is no rush nor pressured incentives.
Current models are not there yet, everyone knows that, but whoever who thinks these models are toys, are not getting the core idea: Models like Astra 6 ARE causal language models, they are bigger and trained longer. It's the same fundamental tech at its core! In fact, even if everything fails, distillation still works! Every single time! Ignoring facts is the only way I can accept that someone would comment that.
Sure, most of these models might look like toys today, but not so much the best ones, and the further you push the line at the scale of the best ones... not to give ideas but those models could be finetuned for good things like targetted research on specific cancer scenarios or really bad automated things that don't require any company's metered safety thresholds.
The smaller you get to do the same things than with the big ones, the more power you have.
OpenAI understood this after releasing GPT-2 and is the reason they no longer open source their models anymore. They would have killed themselves in the process by the way. On the other hand, open source researchers can make incredible discoveries because there is no rush nor pressured incentives.
Don't get me wrong, I don't disagree with you. Of course there is potential, which is partly why some of us are here in the first place.
Also I will bet none of us have OpenAI money.....
It's the same fundamental tech at its core!
Apart from all those techniques that don't work at certain scale (larger or smaller), sure.....
Also GPT-2 was open-weight I thought, not open source? In which case they did release GPT-OSS last year?
My whole point is there was a lot of mudslinging going on regarding model performance as this incident unfolded, and it would do us well to know none of us have OpenAI, Anthropic, etc money and real breakthroughs at this scale are likely to be few and far between.
Lets be real here - most of us are making toy models which are practically useless. It is great as a learning experience and see how far we can push the boundaries with the hardware/money we have to spare. Even if there is something interesting, it may be overshadowed by poor dataset selection or training regimen. Don't see the need for a lot of this childishness I have been observing in some of the previous threads regarding this topic (to say nothing of LLM-as-a-User).
Alright bruh this was just a license request we don’t need to turn evil
In all seriousness yes, they aren’t very useful, but just let us build and create without feeling “useless”
Alright bruh this was just a license request we don’t need to turn evil
In all seriousness yes, they aren’t very useful, but just let us build and create without feeling “useless”
We seem to have extremely different definitions of "evil".
Lemme rephrase what I meant since most people seem to choose to take it as a personal attack instead: "It's not that deep". The whole thing. The license, our work, the copying.
Alright bruh this was just a license request we don’t need to turn evil
In all seriousness yes, they aren’t very useful, but just let us build and create without feeling “useless”
We seem to have extremely different definitions of "evil".
Lemme rephrase what I meant since most people seem to choose to take it as a personal attack instead: "It's not that deep". The whole thing. The license, our work, the copying.
Dude I call everything evil it’s just a joke
I’m not actually mad, and nobody took it as a personal attack. It just seems unnecessary to call it useless. And it’s not like I’m making a demand either. It’s his project, he can do what he wants. I’m just requesting it, as somebody would do if they felt strongly about something.
Dude I call everything evil it’s just a joke
I’m not actually mad, and nobody took it as a personal attack. It just seems unnecessary to call it useless. And it’s not like I’m making a demand either. It’s his project, he can do what he wants. I’m just requesting it, as somebody would do if they felt strongly about something.
We seem to have extremely different definitions of a "joke" as well.
Anyway it was more of a general thought drop since I had been seeing so much discussion regarding this over the past few days. It wasn't directed specifically at you which I should have mentioned before, my bad.
Also you realize I'm putting my own work under that umbrella as well? I prefaced with "practically" because it's just a fact that with current breakthroughs nothing at this scale will be useful as-is. Doesn't mean I think it's a waste of time or a "bad" model, it's just the nature of the constraints we have to deal with.
I wouldn't have cared enough to comment but some of the replies on the previous posts really made me re-evaluate my view of this community, and my participation in it.
Can we all calm down?
Can we all calm down?
yeah
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
I'm so sorry, but these definitely are toy models (I wonder what po$$ibly the one$ that are not could have in common) and a lot of this "real research" would be shredded apart by any reputable peer review process.
It's great to have pride in one's work and push boundaries to their limits. Overselling things is how spaces like this get toxic.
For real--these are hobbyist projects, not real research. None of these "SLM" organizations has released a single paper (to my knowledge), let alone create anything genuinely game-changing.
Man everybody here’s so toxic just let us have fun
Can we all calm down?
No one is angry, were just having a back-and-fourth debate.
Man everybody here’s so toxic just let us have fun
how are people being toxic when some of his points are spot on?
Dude just because there’s no reason to say it. Just have fun. Nobody here is full of themselves. Everybody does it for fun. So what if he wants to make a contribution? Stop acting righteous.
Dude just because there’s no reason to say it. Just have fun. Nobody here is full of themselves. Everybody does it for fun. So what if he wants to make a contribution? Stop acting righteous.
Sorry, I meant some of his points were spot on. I dont think anyone is full of themselves.
Alright bruh this was just a license request we don’t need to turn evil
In all seriousness yes, they aren’t very useful, but just let us build and create without feeling “useless”
We seem to have extremely different definitions of "evil".
Lemme rephrase what I meant since most people seem to choose to take it as a personal attack instead: "It's not that deep". The whole thing. The license, our work, the copying.
It is that deep, though. I mean, personally it is, but in the grand scheme of things, it means absolutely nothing.
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
I'm so sorry, but these definitely are toy models (I wonder what po$$ibly the one$ that are not could have in common) and a lot of this "real research" would be shredded apart by any reputable peer review process.
It's great to have pride in one's work and push boundaries to their limits. Overselling things is how spaces like this get toxic.
For real--these are hobbyist projects, not real research. None of these "SLM" organizations has released a single paper (to my knowledge), let alone create anything genuinely game-changing.
Tell that to my AI girlfriend
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
I'm so sorry, but these definitely are toy models (I wonder what po$$ibly the one$ that are not could have in common) and a lot of this "real research" would be shredded apart by any reputable peer review process.
It's great to have pride in one's work and push boundaries to their limits. Overselling things is how spaces like this get toxic.
For real--these are hobbyist projects, not real research. None of these "SLM" organizations has released a single paper (to my knowledge), let alone create anything genuinely game-changing.
Tell that to my AI girlfriend
NOW THAT is true
These models aren't toy models, most of them (like ACR 1.0 or BananaMind 2.1 Unified) are doing real research here.
I'm so sorry, but these definitely are toy models (I wonder what po$$ibly the one$ that are not could have in common) and a lot of this "real research" would be shredded apart by any reputable peer review process.
It's great to have pride in one's work and push boundaries to their limits. Overselling things is how spaces like this get toxic.
For real--these are hobbyist projects, not real research. None of these "SLM" organizations has released a single paper (to my knowledge), let alone create anything genuinely game-changing.
Tell that to my AI girlfriend
Nahh im scared of her..
Everybody does it for fun.
Exactly my point! Most of us are doing hobbyist work, because they find it fun. But my whole point being I went to the "copied" leaderboard to see what the drama was and saw how people were talking about their "shit models" and "low followers". Also about how that site is now ILLEGAL because of same UI (gasp), while everything was actually according to the license.
If you (general you, not you specifically) call yourself "Labs" and engage in such behavior, you need a reality check regarding the scope of this community.
Everybody does it for fun.
Exactly my point! Most of us are doing hobbyist work, because they find it fun. But my whole point being I went to the "copied" leaderboard to see what the drama was and saw how people were talking about their "shit models" and "low followers". Also about how that site is now ILLEGAL because of same UI (gasp), while everything was actually according to the license.
If you (general you, not you specifically) call yourself "Labs" and engage in such behavior, you need a reality check regarding the scope of this community.
Now THIS has merit. Perhaps I misread you the first time.
Damn didn't expect a licensing argument to get this heated, but def not the worst thing to be passionate abt and very interesting to hear the communities' opinions on this stuff.
Anyways here are my takes:
Take #1
To me, the current license feels less like a rejection of open source principles and more like a response to disappointment, idea/work appropriation, and conflict within the community.
Pretty much nail on the head, as with alot of human society, I believe while broadly legally permissive, there is an expectation of common human decency especially in a small community like this. In the same way when I upload a model under Apache 2.0 (as all my models are), there is the rough expectation under the don't be a dick license that someone isn't gonna copy it weight for weight, upload it as their own model and maliciously make any crediting invisible. And for a long time everyone respected that, however I think this issue has shown that not everyone has the common decency/sense to not do that.
Take #2
If I may pitch in my 2 cents, I fail to see the issue. This leaderboard isn't "software" per se (barring semantics), and I don't see any problem with it having its own license. None of the models are bound to that license and can be used freely. I also like the idea that people can't just copy paste results from here (no issues if they ran the evaluations themselves), which can result in staleness and mis-attributions.
For the most part I agree with this, while we love transformative, innovative, and ethical uses of the code on this leaderboard, as I think I made clear we believe authors have a right to where and how their models are presented, especially on a very important metric like performance, thus we believe we have a responsibility to those who submit their models and evaluations to at the very least add a hurdle to people who may copy and paste these metrics while not respecting this right, especially if the copied location does not have the same tooling that we have built to evaluate the validity of new models.
Anyways that's where I currently stand on this stuff, however this is very much open to change in the future, and in the long run I would like to return it to its original apache 2.0 state. But as it stands currently, I think the current license is suitable.
Dan
(will leave this thread open for further discussion)
Well, that’s not the news I wanted to hear but I respect it. It is your leaderboard, and you can do with it as you want. I appreciate you taking your time to respond, and I understand your reasoning behind the new license.
Whoops sry misclick, meant to press comment
Reopened
The data isn't copyrightable.
Copy metrics for competitor leaderboards or similar products
Makes little sense. OpenAI does exactly this because their Anthropic account was banned.
It's not even legal to do this anyway because the only thing AxiomicLabs owns is the UI for the leaderboard. Everything else is owned by the people who trained the model (if models are copyrightable which it isn't), and the people who wrote the benchmarks.
A computer program feeding texts into a causal LLM and counting correct answers is not human work, and even if it was, it can easily be recreated as it isn't original.
If you want to copyright the UI, that's fine, but using a restrictive license for this serves no purpose. If the UI was generated by AI (it looks like it), it isn't even copyrightable anyway and is public domain no matter what license is attached. Look up "tung tung tung sahur lawsuit" on Google.
It's only now that I discovered the license, so either change it to BSD 2-Clause or Expat, or delete all lines and/or objects in index.html mentioning the strings qikp, Charles, or Kite.
You also can't haphazardly change licenses without notifying the copyright holders. Do that or start requiring a copyright assignment to AxiomicLabs.
The data isn't copyrightable.
Copy metrics for competitor leaderboards or similar products
Makes little sense. OpenAI does exactly this because their Anthropic account was banned.
It's not even legal to do this anyway because the only thing AxiomicLabs owns is the UI for the leaderboard. Everything else is owned by the people who trained the model (if models are copyrightable which it isn't), and the people who wrote the benchmarks.
A computer program feeding texts into a causal LLM and counting correct answers is not human work, and even if it was, it can easily be recreated as it isn't original.
If you want to copyright the UI, that's fine, but using a restrictive license for this serves no purpose. If the UI was generated by AI (it looks like it), it isn't even copyrightable anyway and is public domain no matter what license is attached. Look up "tung tung tung sahur lawsuit" on Google.
Lol UI was created by a combination of AI and I, and ive alr had the discussion abt the data lmao (tldr hand entered data copied from a human writted script more then qualifies for database protection where im from). But by all means feel free to add a removal pr!
e
