🏗️ Building on HF
Adarsh Zolekar
adarshzolekar
AI & ML interests
Exploring AI, ML, Deep Learning, models and datasets while building and contributing to the Hugging Face community.
Recent Activity
updated a dataset 5 days ago
adarshzolekar/foods-nutrition-dataset liked a model 5 days ago
zai-org/GLM-5.3-Flash liked a model 5 days ago
Qwen/Qwen3.8-Flash-NextOrganizations
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 21.3M • 1.48k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 668k • • 975 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 691k • • 5.24k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 823k • • 14.6k
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 253M • • 5.9k -
BAAI/bge-m3
Sentence Similarity • Updated • 37.6M • • 3.5k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 15.7M • 910 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 18.1M • • 1.17k
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.86M • • 6.27k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.76M • • 3.32k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.6M • • 6.86k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.62M • 1.96k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.38M • 2.85k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 12.8M • • 1.37k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.38M • • 3.94k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 12.9M • 991
Reasoning & Agentic Models
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 253M • • 5.9k -
BAAI/bge-m3
Sentence Similarity • Updated • 37.6M • • 3.5k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 15.7M • 910 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 18.1M • • 1.17k
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.86M • • 6.27k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.76M • • 3.32k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.6M • • 6.86k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.62M • 1.96k
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 21.3M • 1.48k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 668k • • 975 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 691k • • 5.24k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 823k • • 14.6k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.38M • 2.85k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 12.8M • • 1.37k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.38M • • 3.94k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 12.9M • 991