Skip to content

ggml-hrx: support BF16 weights - #95

Merged
AaronStGeorge merged 1 commit into
hrx-graph-develop-v2from
users/astgeorg/bf16-llama
Sep 2, 2026
Merged

ggml-hrx: support BF16 weights#95
AaronStGeorge merged 1 commit into
hrx-graph-develop-v2from
users/astgeorg/bf16-llama

Conversation

@AaronStGeorge

@AaronStGeorge AaronStGeorge commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

This PR creates a fairly simple bf16 "dequantization" block and wires it into existing motifs, allowing the existing llama support to now run Llama-3.2-3B-Instruct-BF16.gguf.

perf report generated in ggml-staging-automation (note there are two llama ggufs one is Q4_K and the other bf16 but unfortunately the report rendering doesn't put the quantization somewhere it's easy to read)

@AaronStGeorge
AaronStGeorge merged commit ff2bf7e into hrx-graph-develop-v2 Sep 2, 2026
1 check passed
@AaronStGeorge
AaronStGeorge deleted the users/astgeorg/bf16-llama branch September 2, 2026 00:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants