WHAT’S NEW

New AI models.
Keep up without the noise.

The latest source listings from makers we track, covering language, images, video, audio, and 3D. Catalog checked Sep 10, 2026.

Sorted by the date a source repository was created, which may precede public release. This is a feed of new listings, not confirmed release announcements. Downloadable weights do not automatically mean an open-source license. Coverage and date meanings.

Audio listings

62 dated records

Qwen · Alibaba · Downloadable

Qwen3-ASR-1.7B

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 28, 2026
Parameters
1.7B
4-bit weights only
≈ 0.9 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-ASR-0.6B

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 28, 2026
Parameters
600M
4-bit weights only
≈ 0.3 GB
License
apache-2.0
Compare this model →

Mistral AI · Downloadable

Voxtral-Mini-4B-Realtime-2602

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 21, 2026
Parameters
4B
4-bit weights only
≈ 2 GB
License
apache-2.0
Compare this model →

Microsoft · Downloadable

VibeVoice-ASR

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 21, 2026
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-TTS-12Hz-0.6B-Base

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 21, 2026
Parameters
600M
4-bit weights only
≈ 0.3 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-TTS-Tokenizer-12Hz

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 21, 2026
Parameters
Not reported
4-bit weights only
Not reported
License
apache-2.0
Compare this model →

Microsoft · Downloadable

paza-whisper-large-v3-turbo

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 8, 2026
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Microsoft · Downloadable

paza-Phi-4-multimodal-instruct

VisionTranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 7, 2026
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Google · Downloadable

medasr

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Dec 18, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
other
Compare this model →

Z.ai · Downloadable

GLM-TTS

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Dec 10, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Z.ai · Downloadable

GLM-ASR-Nano-2512

TranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Dec 9, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Microsoft · Downloadable

VibeVoice-Realtime-0.5B

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Dec 4, 2025
Parameters
500M
4-bit weights only
≈ 0.3 GB
License
mit
Compare this model →

Mistral AI · Downloadable

Voxtral-4B-TTS-2603

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Nov 17, 2025
Parameters
4B
4-bit weights only
≈ 2 GB
License
cc-by-nc-4.0
Compare this model →

Microsoft · Downloadable

VibeVoice-1.5B

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Aug 25, 2025
Parameters
1.5B
4-bit weights only
≈ 0.8 GB
License
mit
Compare this model →

Mistral AI · Downloadable

Voxtral-Small-24B-2507

Capabilities unconfirmed

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jul 1, 2025
Parameters
24B
4-bit weights only
≈ 12 GB
License
apache-2.0
Compare this model →

Stability AI · Downloadable

stable-audio-open-small

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
May 12, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
other
Compare this model →

Moonshot AI · Downloadable

Kimi-Audio-7B

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Apr 25, 2025
Parameters
7B
4-bit weights only
≈ 3.5 GB
License
mit
Compare this model →

Moonshot AI · Downloadable

Kimi-Audio-7B-Instruct

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Apr 25, 2025
Parameters
7B
4-bit weights only
≈ 3.5 GB
License
mit
Compare this model →

Microsoft · Downloadable

Phi-4-multimodal-instruct-onnx

VisionTranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Feb 25, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Microsoft · Downloadable

Phi-4-multimodal-instruct

VisionTranscriptionAudio understanding

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Feb 24, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
mit
Compare this model →

Stability AI · Downloadable

stable-codec-speech-16k-base

Audio generation

Works with speech or audio. Check the model card for transcription and generation capabilities.

Source listed
Jan 14, 2025
Parameters
Not reported
4-bit weights only
Not reported
License
other
Compare this model →

What makes this feed useful?

Capability tags distinguish a model that reads images from one that creates them. Each listing links to its source and a permanent profile, so you can inspect its license and compare it with a model you already know.

Does this cover every new release?

No. The current feed tracks public Hugging Face metadata from selected makers. It can miss untracked makers, cloud-only releases, and records beyond a source’s retrieval limit. Announced models without a source listing are not included. Check Sources & methodology for the current coverage.

Listings update with the catalog automatically. A missing or delayed source can leave older information visible; each model profile shows when it was checked.