🏗️ Building on HF
Adarsh Zolekar
adarshzolekar
AI & ML interests
Exploring AI, ML, Deep Learning, models and datasets while building and contributing to the Hugging Face community.
Recent Activity
liked a model about 18 hours ago
google/embeddinggemma-2 upvoted a paper about 18 hours ago
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations upvoted a paper about 18 hours ago
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision ReliabilityOrganizations
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 19.5M • 1.58k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 175k • • 983 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 631k • • 5.43k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 662k • • 15.5k
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 226M • • 6.24k -
BAAI/bge-m3
Sentence Similarity • Updated • 33M • • 3.84k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 12.9M • 950 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 16.4M • • 1.23k
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 3.79M • • 6.59k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.32M • • 3.42k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 10.8M • • 7.19k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.25M • 2.1k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.08M • 3.74k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 9.82M • • 2.08k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.49M • • 4.02k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 5.13M • 1.13k
Reasoning & Agentic Models
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 226M • • 6.24k -
BAAI/bge-m3
Sentence Similarity • Updated • 33M • • 3.84k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 12.9M • 950 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 16.4M • • 1.23k
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 3.79M • • 6.59k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.32M • • 3.42k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 10.8M • • 7.19k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.25M • 2.1k
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 19.5M • 1.58k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 175k • • 983 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 631k • • 5.43k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 662k • • 15.5k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.08M • 3.74k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 9.82M • • 2.08k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.49M • • 4.02k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 5.13M • 1.13k