Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German Paper • 2406.06131 • Published Jun 10, 2024
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Paper • 2506.06275 • Published Jun 6, 2025
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Paper • 2606.06177 • Published Jun 4
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese Paper • 2603.26511 • Published Mar 27
Position: Hippocampal Explicit Memory Is the Cornerstone for AGI Paper • 2606.11245 • Published Jun 5
MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control Paper • 2407.02736 • Published Jul 3, 2024
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture Paper • 2310.03052 • Published Oct 4, 2023 • 4
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources Paper • 2509.25531 • Published Sep 29, 2025 • 11
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution Paper • 2510.08697 • Published Oct 9, 2025 • 41
BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications Paper • 2509.24908 • Published Sep 29, 2025 • 3
EmbeddingGemma: Powerful and Lightweight Text Representations Paper • 2509.20354 • Published Sep 24, 2025 • 51
What Language Model to Train if You Have One Million GPU Hours? Paper • 2210.15424 • Published Oct 27, 2022 • 2
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Paper • 2211.05100 • Published Nov 9, 2022 • 40
Enhancing Few-shot Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design Strategies Paper • 2305.12586 • Published May 21, 2023