AI & ML interests

None defined yet.

AtAndDevย 
posted an update about 1 month ago
view post
Post
2886
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M

Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).

Thanks for everyone supporting!
  • 3 replies
ยท
AtAndDevย 
posted an update about 1 month ago
view post
Post
208
SPECK1.5 IS COMING SOON!
Same 5B token budget but much better corpus quality.

Also getting a ton of downloads, thanks for everyone downloading and liking <3

specklabs
AtAndDevย 
posted an update about 1 month ago
view post
Post
160
NEW SPECK UPDATES:

Just hit #14 and #15 with out FIRST models on Open SLM Leaderboard. The models were trained on 5B tokens, while competing with similarly sized models trained on more than 6-20x the data.

A new base model Speck1.5-140M being trained right now on a higher quality corpus and will be released soon.
SpeckChat3 is coming very soon with 1 million samples, specifically designed to post train small base models.

Also, just to clarify stuff, we will NOT release anything that is NOT MIT licensed EVER. Openness is needed in small language research.

Thanks to everyone supporting the project, and stay tuned for new releases!
AtAndDevย 
posted an update about 2 months ago
view post
Post
1894
SPECK UPDATES:
1 New instruct model tuned on top of Speck1-140M: specklabs/Speck1-140M-Instruct
2 Instruction tuning datasets
2 GGUFs

Much more coming soon:
Speck1.1-140M-Instruct that is post trained on SpeckChat2 will be coming very soon
New base model Speck1.5-140M is coming with a much higher quality corpus

Thanks to everyone who is already supporting the project, and stay tuned for new releases!
  • 3 replies
ยท
AtAndDevย 
posted an update about 2 months ago
view post
Post
2130
FIRST SPECK MODEL RELEASED:
specklabs/Speck1-140M

new models coming very soon (both instruct and much better models), with much much higher training scale as i am getting marenostrum5 access soon!
we will be looking at 100b-2t token budgets :)
  • 4 replies
ยท
s3nhย 
posted an update 3 months ago
view post
Post
526
Uncensoring Mistral,

give it a try
s3nh/Ministral-3-14B-Instruct-2512-BF16-abliterated
s3nhย 
posted an update 4 months ago
view post
Post
563
Existing methods โ€” GPTQ, AWQ, llama.cpp's k-quants โ€” minimize empirical loss heuristically. None of them prove they are optimal in any information-theoretic sense. ICRB-Q builds a quantization scheme that is provably optimal via the Cramรฉr-Rao lower bound (CRB): no unbiased estimator of a weight can have lower variance than [F(ฮธ)]โปยน, where F is the Fisher information matrix.
  • 1 reply
ยท
s3nhย 
posted an update 12 months ago
view post
Post
950
Eduhelp with more empathy, based on model finetuned on
psychotheraputic preferences just landed on


Beck-8B as a base model, 13000 steps on educational dataset.
Time to go further and build more ๐Ÿฅฐ
s3nh/EduHelp_Beck_8B
Thanks to @basilic_ai for computations <3
s3nhย 
posted an update 12 months ago
view post
Post
4432
Just tried to create an educational assistant for younger people who can struggle with visualsation of 'what is this sorcery all about'.
Its first step of my spare time projects, sft on Qwen3-8B,

EduHelper is a child-friendly tutoring assistant fine-tuned from the Qwen3-8B base model using parameter-efficient fine-tuning (PEFT) with LoRA on the ajibawa-2023/Education-Young-Children dataset.

s3nh/EduHelp-8B

Glad to share my work, have a wonderful day!
  • 2 replies
ยท
AtAndDevย 
posted an update about 1 year ago
view post
Post
786
Qwen 3 Coder is a personal attack to k2, and I love it.
It achieves near SOTA on LCB while not having reasoning.
Finally people are understanding that reasoning isnt necessary for high benches...

Qwen ftw!

DECENTRALIZE DECENTRALIZE DECENTRALIZE
AtAndDevย 
posted an update over 1 year ago
view post
Post
3258
deepseek-ai/DeepSeek-R1-0528

This is the end
  • 2 replies
ยท
AtAndDevย 
posted an update over 1 year ago
view post
Post
3164
Llama 4 is out...
  • 3 replies
ยท
AtAndDevย 
posted an update over 1 year ago
view post
Post
4400
There seems to multiple paid apps shared here that are based on models on hf, but some ppl sell their wrappers as "products" and promote them here. For a long time, hf was the best and only platform to do oss model stuff but with the recent AI website builders anyone can create a product (really crappy ones btw) and try to sell it with no contribution to oss stuff. Please dont do this, or try finetuning the models you use...
Sorry for filling yall feed with this bs but yk...
  • 6 replies
ยท
AtAndDevย 
posted an update over 1 year ago
view post
Post
1688
Gemma 3 seems to be really good at human preference. Just waiting for ppl to see it.
AtAndDevย 
posted an update over 1 year ago
view post
Post
2513
@nroggendorff is that you sama?
  • 2 replies
ยท
AtAndDevย 
posted an update over 1 year ago
view post
Post
1962
everywhere i go i see his face
AtAndDevย 
posted an update over 1 year ago
view post
Post
597
Deepseek gang on fire fr fr
AtAndDevย 
posted an update over 1 year ago
view post
Post
1673
R1 is out! And with a lot of other R1 releated models...
s3nhย 
posted an update almost 2 years ago
view post
Post
2556
Welcome back,

Small Language Models Enthusiasts and GPU Poor oss enjoyers lets connect.
Just created an organization which main target is to have fun with smaller models tuneable on consumer range GPUs, feel free to join and lets have some fun, much love ;3

SmolTuners
  • 3 replies
ยท
AtAndDevย 
posted an update almost 2 years ago
view post
Post
510
@s3nh Hey man check your discord! Got some news.
  • 4 replies
ยท