I ran a similar experiment last month and ported Qwen 3 Omni to llama cpp. I was able to get GGUF conversion, quantization, and all input and output modalities working in less than a week. I submitted the work as a PR to the codebase and understandably, it was rejected.https://github.com/ggml-org/llama.cpp/pull/18404
https://huggingface.co/TrevorJS/Qwen3-Omni-30B-A3B-GGUF
antirez|1 month ago
rjh29|1 month ago
nickandbro|1 month ago
nickpsecurity|1 month ago
[deleted]
unknown|1 month ago
[deleted]