Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
AI & ML interests
Factuality, reasoning, alignment, LLM applications
Recent Activity
View all activity
Papers
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
spaces 7
Running
LudoBench
🎲
Multimodal Game Reasoning Benchmark [ICLR 2026]
Sleeping
Agents
Answer Convergence Early Stopping
🛑
Demo for EMNLP Paper "Answer Convergence as a Signal..."
Sleeping
FactRBench
🏆
View and analyze long-form factuality leaderboard
Running
3
ExpertLongBench
🚀
Leaderboard for ExpertLongBench
Sleeping
1
ManyICLBench
🚀
Leaderboard for ManyICLBench
Running
MLRC-BENCH
📊
Display model performance rankings
models 15
launch/MET-D-Gemma3-4B-en-only
Text Generation • 4B • Updated • 32
launch/MET-D-Gemma3-4B
Text Generation • 4B • Updated • 32
launch/MET-D-Qwen3-8B-en-only
Text Generation • 8B • Updated • 20
launch/MET-D-Qwen3-8B
Text Generation • 8B • Updated • 34
launch/MET-D-Qwen3-4B-zh-only
Text Generation • 4B • Updated • 35
launch/MET-D-Qwen3-4B-ms-only
Text Generation • 4B • Updated • 34
launch/MET-D-Qwen3-4B-ko-only
Text Generation • 4B • Updated • 33
launch/MET-D-Qwen3-4B-hi-only
Text Generation • 4B • Updated • 17
launch/MET-D-Qwen3-4B-es-only
Text Generation • 4B • Updated • 35
launch/MET-D-Qwen3-4B-en-only
Text Generation • 4B • Updated • 34
datasets 14
launch/MCLASH
Viewer • Updated • 2.61k • 119
launch/CLASH
Viewer • Updated • 345 • 199 • 3
launch/thinkprm-1K-verification-cots
Viewer • Updated • 1k • 62 • 8
launch/LudoBench
Viewer • Updated • 638 • 62
launch/ExpertLongBench
Preview • Updated • 360 • 10
launch/ManyICLBench
Viewer • Updated • 66 • 150 • 1
launch/CMV
Viewer • Updated • 133 • 27
launch/FactRBench
Viewer • Updated • 1.06k • 15 • 2
launch/FactBench
Viewer • Updated • 1k • 92 • 3
launch/gov_report
Viewer • Updated • 58.4k • 1.7k • 14