Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions".
AI & ML interests
LLM reasoning
Recent Activity
View all activity
models 7
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch3
Text Generation • 9B • Updated • 273
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch2
Text Generation • 9B • Updated • 300
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch1
Text Generation • 9B • Updated • 302
Miaow-Lab/RLVR-Linearity-Checkpoints
Text Generation • Updated
Miaow-Lab/STT-Agent-RL
196k • Updated • 20 • 1
Miaow-Lab/STT-Agent-SFT
196k • Updated • 9 • 1
Miaow-Lab/SSAE-Checkpoints
Feature Extraction • Updated
datasets 6
Miaow-Lab/v7-kernel-candidates
Viewer • Updated • 5.16k • 58
Miaow-Lab/STT-Arena
Preview • Updated • 60 • 2
Miaow-Lab/OpenSkillRisk
Updated • 88
Miaow-Lab/RUT-Bench
Viewer • Updated • 1.64k • 30 • 1
Miaow-Lab/RLVR-Linearity-Dataset
Viewer • Updated • 40.3k • 40
Miaow-Lab/SSAE-Dataset
Viewer • Updated • 1.28M • 30