BeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 2 months agounsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coexternal-linkmessage-square8linkfedilinkarrow-up126arrow-down11file-text
arrow-up125arrow-down1external-linkunsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coBeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 2 months agomessage-square8linkfedilinkfile-text
minus-squareMultiplexer@discuss.tchncs.delinkfedilinkEnglisharrow-up3·2 months agoAnyone knows, why 4bit quant is only marginally smaller than 8bit quant, though? And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
minus-squareBeefAndPoultry@lemmus.orgOPlinkfedilinkEnglisharrow-up4·2 months agothe original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher
Anyone knows, why 4bit quant is only marginally smaller than 8bit quant, though?
And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
the original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher