llama.cpp

mirror of https://github.com/ggerganov/llama.cpp synced 2026-04-26 13:32:04 +02:00

History

Stephen Cox 547765a93e mtmd: add Gemma 4 audio conformer encoder support (#21421 ) * mtmd: add Gemma 4 audio conformer encoder support Add audio processing for Gemma 4 E2B/E4B via a USM-style Conformer. Architecture: - 12-layer Conformer: FFN → Self-Attention → Causal Conv1D → FFN → Norm - Subsampling Conv Projection: 2x Conv2D(stride=2) with LayerNorm - Full self-attention with sinusoidal RPE and sliding window mask (24) - Logit softcapping at 50.0, ClippableLinear clamping - Output: 1024 → 1536 → RMSNorm → multimodal embedder Mel preprocessing (dedicated mtmd_audio_preprocessor_gemma4a): - HTK mel scale, 128 bins, magnitude STFT, mel_floor=1e-3 - Standard periodic Hann window (320 samples), zero-padded to FFT size - Semicausal left-padding (frame_length/2 samples) - Frame count matched to PyTorch (unfold formula) - No pre-emphasis, no Whisper-style normalization - Mel cosine similarity vs PyTorch: 0.9998 Key fixes: - Tensor loading dedup: prevent get_tensor() from creating duplicate entries in ctx_data. Fixed with std::set guard. - ClippableLinear clamp_info loading moved after per-layer tensors. - Sliding window mask (24 positions) matching PyTorch context_size. - Skip Whisper normalization for Gemma4 mel output. Tested on E2B and E4B with CPU and Vulkan backends. Transcribes: "Glad to see things are going well and business is starting to pick up" (matching ground truth). Ref: #21325		2026-04-12 14:15:26 +02:00
..
batched-bench	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
cli	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
completion	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
cvector-generator	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
export-lora	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
fit-params	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
gguf-split	gguf-split : clarify operation of gguf-split (#19749 )	2026-03-25 13:12:50 +02:00
imatrix	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
llama-bench	common : add callback interface for download progress (#21735 )	2026-04-10 22:17:00 +02:00
mtmd	mtmd: add Gemma 4 audio conformer encoder support (#21421 )	2026-04-12 14:15:26 +02:00
parser	common/parser: fix call ID detection (Mistral parser mostly) + atomicity for tag-json parsers (#21230 )	2026-04-03 17:51:52 +02:00
perplexity	ggml: backend-agnostic tensor parallelism (experimental) (#19378 )	2026-04-09 16:42:19 +02:00
quantize	ggml: add Q1_0 1-bit quantization support (CPU) (#21273 )	2026-04-06 20:55:21 +02:00
results	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
rpc	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
server	fix: Proper messages rendering for "Show raw output" (#21672 )	2026-04-12 13:08:11 +02:00
tokenize	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
tts	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
CMakeLists.txt	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00