I may or may not have splurged on a 128GB AMD Ryzen AI Max 395 (strix halo) system for ‘AI stuff’ (told you I was a noob).

I’ve been running Ubuntu on it with the AMD drivers (think its ROCm?), ollama and seems to be working fine.

An LLM told me to change the RAM/VRAM ratio to 50:50 (so 64GB for the CPU, 64GB for the GPU). I dunno if that was correct, seems like a waste tbh. Feels like I could give the GPU more resources and run bigger models.

I’ve read about Lemonade being better than Ollama on strix halo? Also, I realised that I might not be using the NPU as extra work is required to get that up and running.

I’m looking for advice from users on the same hardware. What OS are you using? How do you have the RAM/VRAM ratio configured? What’s your stack? That sorta thing.

PS - If my wife asks, the machine only cost like £250 and is a second-hand floor model.

  • Luminous5481 "Enemy of the State"@anarchist.nexus
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    7 days ago

    check out HaloFPX. I can dig up the URL if you can’t find it. it’s a runtime that’s made for Strix Halo that ships with custom tuned models for ROCmFPX (I think that’s what it’s called), and it can give you really good speeds.

    one of the models it comes with is a custom Ornith 1.5 35B, and with HaloFPX I get speeds in Deepseek Harness that are so fast I can’t read fast enough to keep up with generation.

    EDIT link: https://github.com/julianmb/halofpx