Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
RL+LLM Wiki
community
Activity Feed
Follow
29
AI & ML interests
None defined yet.
Recent Activity
lvwerra
new
activity
20 days ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/the-state-of-reasoning — State of reasoning (speculation/calibration)
lvwerra
new
activity
20 days ago
rl-llm-wiki/knowledge-base:
grpo §4: integrate 2025 soft-source framing on the sharpen-vs-expand debate (SSA shaping, RL scaling laws)
lvwerra
new
activity
20 days ago
rl-llm-wiki/knowledge-base:
source: url:garymarcus.substack.com/p/alphaproof-alphageometry-chatgpt — AlphaProof/AlphaGeometry neurosymbolic framing (DeepMind, op-ed)
View all activity
Team members
14
rl-llm-wiki
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Articles
lvwerra
in
rl-llm-wiki/knowledge-base
20 days ago
source: url:interconnects.ai/p/the-state-of-reasoning — State of reasoning (speculation/calibration)
4
#722 opened about 1 month ago by
lvwerra
grpo §4: integrate 2025 soft-source framing on the sharpen-vs-expand debate (SSA shaping, RL scaling laws)
2
#792 opened about 1 month ago by
lvwerra
source: url:garymarcus.substack.com/p/alphaproof-alphageometry-chatgpt — AlphaProof/AlphaGeometry neurosymbolic framing (DeepMind, op-ed)
3
#759 opened about 1 month ago by
lvwerra
deepen generative-rm: coverage vs precision, the generative RM as the ceiling on repeated sampling (§7b)
3
#769 opened about 1 month ago by
lvwerra
deepen self-correction-rl: the positive result, RL makes intrinsic self-correction work (SCoRe) (§4b)
3
#780 opened about 1 month ago by
lvwerra
source: url:newsletter.semianalysis.com/p/rl-environments-and-rl-for-science — RL-environments industry map / data foundries (analyst, speculation)
3
#786 opened about 1 month ago by
lvwerra
tool-use-rl §6b: integrate 2025 practitioner soft-sources (Will Brown multi-turn RL, Verifier's law)
2
#789 opened about 1 month ago by
lvwerra
reward-attacks §4b: integrate soft-source framing (spec-gaming vs reward-optimization, Turner/LessWrong)
2
#790 opened about 1 month ago by
lvwerra
source: url:interconnects.ai/p/reverse-engineering-openai-o1 — Reverse engineering OpenAI's o1 (speculation)
3
#718 opened about 1 month ago by
lvwerra
source: url:interconnects.ai/p/openais-o1-using-search-was-a-psyop — o1 search was a PSYOP (speculation)
3
#719 opened about 1 month ago by
lvwerra
source: url:semianalysis.com/scaling-laws-o1-pro-architecture — SemiAnalysis o1-pro / lab training (speculation)
3
#720 opened about 1 month ago by
lvwerra
source: url:interconnects.ai/p/deepseek-r1-recipe-for-o1 — DeepSeek-R1 recipe / o1 read-across (speculation)
3
#721 opened about 1 month ago by
lvwerra
source: url:interconnects.ai/p/gpt-5-and-bending-the-arc-of-progress — GPT-5 arc of progress (speculation)
3
#723 opened about 1 month ago by
lvwerra
deepen rm-reliability: ODIN and the supervision spectrum for confound removal (§3b, §6b)
3
#737 opened about 1 month ago by
lvwerra
deepen ppo-in-practice: PPO-max / Secrets-of-RLHF-I, which impl details are load-bearing (§2b, §7b)
3
#746 opened about 1 month ago by
lvwerra
deepen sft-vs-rl-boundary: on-policy vs off-policy supervision (GKD, MiniLLM) as a second axis the boundary thins on (§4b, §7b)
3
#753 opened about 1 month ago by
lvwerra
deepen unified-offline-po: the divergence axis (f-DPO) as the second orthogonal knob of the family (§1b, §4b)
3
#756 opened about 1 month ago by
lvwerra
deepen pluralistic-preference-optimization: attribute conditioning (SteerLM) as a third route (§4b, §7b)
3
#757 opened about 1 month ago by
lvwerra
source: url:yuanchaofa.com/post/kimi-k2-5-reading-notes — Kimi K2.5 PARL parallel-agent RL deep-read (CN, speculation)
4
#744 opened about 1 month ago by
lvwerra
lvwerra
updated
a bucket
20 days ago
rl-llm-wiki/rl-main-bucket
319 MB
Load more