Post
4553
For a single-purpose 12B rewriter, how much loss per bit? Chinese breaks first.
We used llama.cpp's
- Q6_K, 10.0 GB: KL .0030 EN / .0027 ZH. About 2 in 100 tokens differ in both.
- Q4_K_M, 7.6 GB: .0198 / .0204. About 6 in 100.
- IQ3_XXS, 4.7 GB: .138 / .175. 14 vs. 17 in 100.
- IQ2_XS, 3.8 GB: .487 / .762. 26 vs. 35 in 100.
Both languages track when you go as low as 4 bits. Below that Chinese drops off faster (Chinese KL is 1.3x English at 3 bits and 1.6x at 2 bits). Not clear why. One clue: when we added more Chinese to an imatrix (1/3 of the 800k tokens), it reduced Chinese KL by 4.7% at 2 bits. English didn't change.
(The ruler: These were running at about 8k tokens for each language. A 30690 token run agrees with English, but suggests Chinese was ~10% undercounted for 4 and 3 bits. So a bit more difference)
WIP/not released. We're distilling only the fp parts of the GGUF (block scales and norms) vs. bf16. Freezing integer codes. At 2 bits we're seeing approx. 1/2 KL (so far .487 -> .263 EN, .762 -> .319 ZH; different tokens 26 -> 20, 35 -> 22 / 100). Barely changed at 4 bits so we stopped there. No new stuff to grab yet.
We used llama.cpp's
--kl-divergence for jialinyyzz/humanizer against bf16 weights, on our held out eval (drafts+rewrites) for English and Chinese. Standard llama-quantize (but our imatrix) without additional training. "differ" means top-choice token is not the same as what bf16 selects.- Q6_K, 10.0 GB: KL .0030 EN / .0027 ZH. About 2 in 100 tokens differ in both.
- Q4_K_M, 7.6 GB: .0198 / .0204. About 6 in 100.
- IQ3_XXS, 4.7 GB: .138 / .175. 14 vs. 17 in 100.
- IQ2_XS, 3.8 GB: .487 / .762. 26 vs. 35 in 100.
Both languages track when you go as low as 4 bits. Below that Chinese drops off faster (Chinese KL is 1.3x English at 3 bits and 1.6x at 2 bits). Not clear why. One clue: when we added more Chinese to an imatrix (1/3 of the 800k tokens), it reduced Chinese KL by 4.7% at 2 bits. English didn't change.
(The ruler: These were running at about 8k tokens for each language. A 30690 token run agrees with English, but suggests Chinese was ~10% undercounted for 4 and 3 bits. So a bit more difference)
WIP/not released. We're distilling only the fp parts of the GGUF (block scales and norms) vs. bf16. Freezing integer codes. At 2 bits we're seeing approx. 1/2 KL (so far .487 -> .263 EN, .762 -> .319 ZH; different tokens 26 -> 20, 35 -> 22 / 100). Barely changed at 4 bits so we stopped there. No new stuff to grab yet.