Specialized and Edge Language Models (GLM, Hunyuan, Granite, LFM, LLM-jp, etc.)
i love SSD
aoiandroid
AI & ML interests
ML Engineer specializing in Generative AI & fine-tuning across text, translation, image, and video synthesis. Building applied multimodal models. now interested in ANE optimization.
Recent Activity
updated a model 3 days ago
aoiandroid/whisper-ur-large-v3-turbo-whisperkit-coreml-macos published a model 3 days ago
aoiandroid/whisper-ur-large-v3-turbo-whisperkit-coreml-macos updated a model 3 days ago
aoiandroid/whisper-ur-large-v3-turbo-whisperkit-coreml-iosOrganizations
Audio, VAD & Diarization
Voice Activity Detection, Diarization, and Audio Processing models
Nemotron
NVIDIA Nemotron ASR and VoiceChat models
Papers
-
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Paper • 2509.24832 • Published -
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 87 -
Map2World: Segment Map Conditioned Text to 3D World Generation
Paper • 2605.00781 • Published • 25 -
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
Paper • 2604.24026 • Published • 22
mms-lid
translategemma
parakeet
madlad
whisper
-
aoiandroid/whisper-distil-whisper-distil-large-v3-coreml
Updated • 17 -
aoiandroid/whisper-distil-whisper-distil-large-v3-594MB-coreml
Updated • 3 -
aoiandroid/whisper-distil-whisper-distil-large-v3-turbo-coreml
Updated • 10 -
aoiandroid/whisper-distil-whisper-distil-large-v3-turbo-600MB-coreml
Updated • 2
omnivoice
OCR & Document AI
OCR, Document Parsing, and Multimodal Vision models
Kokoro & Voice TTS
Kokoro, VibeVoice, CosyVoice, and Text-to-Speech / Voice generation models
Game Dev
neucodec
aoiandroid models matching neucodec search
yolo
gemma
llama
aoiandroid models matching llama search
nllb
qwen
neutts-jp
Other LLM / SLM
Specialized and Edge Language Models (GLM, Hunyuan, Granite, LFM, LLM-jp, etc.)
OCR & Document AI
OCR, Document Parsing, and Multimodal Vision models
Audio, VAD & Diarization
Voice Activity Detection, Diarization, and Audio Processing models
Kokoro & Voice TTS
Kokoro, VibeVoice, CosyVoice, and Text-to-Speech / Voice generation models
Nemotron
NVIDIA Nemotron ASR and VoiceChat models
Game Dev
Papers
-
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Paper • 2509.24832 • Published -
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 87 -
Map2World: Segment Map Conditioned Text to 3D World Generation
Paper • 2605.00781 • Published • 25 -
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
Paper • 2604.24026 • Published • 22
neucodec
aoiandroid models matching neucodec search
mms-lid
yolo
translategemma
gemma
parakeet
llama
aoiandroid models matching llama search
madlad
nllb
whisper
-
aoiandroid/whisper-distil-whisper-distil-large-v3-coreml
Updated • 17 -
aoiandroid/whisper-distil-whisper-distil-large-v3-594MB-coreml
Updated • 3 -
aoiandroid/whisper-distil-whisper-distil-large-v3-turbo-coreml
Updated • 10 -
aoiandroid/whisper-distil-whisper-distil-large-v3-turbo-600MB-coreml
Updated • 2
qwen
omnivoice
neutts-jp