first commit

This commit is contained in:
Donato Capitella
2025-11-30 08:06:04 +00:00
commit 82c6e04253
34 changed files with 50959 additions and 0 deletions
+666
View File
@@ -0,0 +1,666 @@
:root {
--bg: #f5f6fa;
--ink: #101828;
--muted: #6b7080;
--accent: #155eef;
--border: #d8dce6;
--card: #ffffff;
--chip-bg: #e6ecff;
--chip-active-bg: #155eef;
--chip-active-ink: #fff;
--winner-bg: #d7f5e3;
--winner-ink: #025333;
--warn: #c2410c;
--model-col: 180px;
--winner-col: 120px;
}
* {
box-sizing: border-box;
}
body {
margin: 0;
font: 13px/1.35 "Inter", "Segoe UI", system-ui, -apple-system, sans-serif;
background: var(--bg);
color: var(--ink);
}
header {
padding: 14px 20px 4px;
background: var(--card);
border-bottom: 1px solid var(--border);
}
header h1 {
margin: 0 0 4px;
font-size: 20px;
font-weight: 600;
}
header p {
margin: 2px 0;
font-size: 12px;
color: var(--muted);
}
.controls,
.panel {
background: var(--card);
border-bottom: 1px solid var(--border);
padding: 10px 20px;
}
.controls {
display: flex;
gap: 12px;
flex-wrap: wrap;
align-items: flex-start;
}
.control {
min-width: 200px;
}
.control.grow {
flex: 1 1 320px;
}
.slider-block {
min-width: 260px;
}
label {
display: block;
font-size: 10px;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--muted);
margin-bottom: 3px;
}
input[type="text"],
select {
width: 100%;
padding: 6px 9px;
border-radius: 6px;
border: 1px solid var(--border);
font-size: 13px;
background: #fff;
}
.chip-row {
display: flex;
flex-wrap: wrap;
gap: 4px;
}
.chip {
border: none;
border-radius: 999px;
padding: 3px 10px;
font-size: 12px;
cursor: pointer;
background: var(--chip-bg);
color: var(--ink);
}
.chip.active {
background: var(--chip-active-bg);
color: var(--chip-active-ink);
}
.chip.small {
font-size: 11px;
padding: 3px 8px;
}
.panel.compact {
padding: 8px 20px;
}
.panel-split {
display: flex;
gap: 16px;
flex-wrap: wrap;
align-items: center;
}
.backend-list {
display: flex;
flex-wrap: wrap;
gap: 6px 14px;
}
.backend-label {
display: flex;
align-items: center;
gap: 8px;
}
.backend-actions {
display: flex;
gap: 6px;
}
.backend-item {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 12px;
color: var(--ink);
text-transform: none;
}
.backend-item input {
transform: translateY(1px);
}
.backend-item .tag {
font-size: 10px;
padding: 0 6px;
border-radius: 999px;
background: #eef2ff;
color: #1d3ea5;
transform: translateY(-2px);
}
.backend-item .tag.tag-hblt0 {
background: #e9edff;
color: #1d3ea5;
}
.backend-item .tag.tag-rocwmma {
background: #eef9ff;
color: #0a517a;
}
.backend-item .tag.tag-rocwmma-improved {
background: #faf3ff;
color: #6b1fb7;
}
.backend-item .tag.tag-improved {
background: #fef9e7;
color: #8a5a00;
}
.stats-box {
margin-left: auto;
display: flex;
gap: 10px;
align-items: center;
font-size: 12px;
color: var(--muted);
}
#tables {
display: grid;
gap: 14px;
}
.test-block h2 {
margin: 0 0 4px;
font-size: 12px;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--muted);
}
.table-wrap {
border-radius: 8px;
border: 1px solid var(--border);
background: var(--card);
position: relative;
width: 100%;
max-width: 100%;
overflow: hidden;
}
.table-scroll {
overflow-x: auto;
overflow-y: hidden;
width: 100%;
position: relative;
scrollbar-gutter: stable both-edges;
display: block;
}
.table-scroll table {
min-width: 100%;
}
table {
border-collapse: collapse;
font-size: 11.5px;
width: max-content;
min-width: 100%;
table-layout: fixed;
}
thead {
background: #f4f6fb;
}
th,
td {
padding: 4px 6px;
border-bottom: 1px solid var(--border);
white-space: normal;
border-right: 1px solid var(--border);
overflow-wrap: anywhere;
}
th {
position: relative;
font-weight: 600;
}
th.sticky,
td.sticky {
position: sticky;
left: 0;
background: inherit;
z-index: 3;
box-shadow: 1px 0 0 var(--border);
}
th.model,
td.model {
width: var(--model-col);
position: sticky;
left: 0;
z-index: 3;
background: #f8f9ff;
}
th.winner,
td.winner {
width: var(--winner-col);
position: sticky;
left: var(--model-col);
z-index: 3;
background: #f1f5ff;
}
td.model {
min-width: 170px;
font-weight: 500;
}
td.model .model-head {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 6px;
}
.model-pill {
display: inline-flex;
align-items: center;
padding: 2px 8px;
border-radius: 999px;
font-size: 10px;
text-transform: uppercase;
letter-spacing: 0.05em;
background: #eceff5;
color: #27303f;
border: 1px solid transparent;
}
.model-pill-rpc {
background: #fdf2f8;
border-color: #fbcfe8;
color: #9d174d;
}
.model-pill-rocwmma {
background: #eef9ff;
border-color: #c7e9ff;
color: #0a517a;
}
.legend {
display: flex;
flex-direction: column;
gap: 6px;
margin-top: 8px;
}
.legend label {
font-size: 10px;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--muted);
}
.legend-pills {
display: flex;
flex-wrap: wrap;
gap: 8px;
}
.legend-pill {
display: inline-flex;
align-items: center;
gap: 4px;
border-radius: 999px;
border: 1px solid transparent;
background: #e9edff;
color: var(--ink);
}
.legend-pill-default {
background: #e9edff;
color: var(--ink);
}
.legend-pill-rpc {
background: #fdf2f8;
border-color: #fbcfe8;
color: #9d174d;
}
.legend-pill-rocwmma {
background: #eef9ff;
border-color: #c7e9ff;
color: #0a517a;
}
.legend-pill-rocwmma-improved {
background: #faf3ff;
border-color: #e0c8ff;
color: #6b1fb7;
}
.modal.hidden {
display: none;
}
.modal {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.5);
display: flex;
align-items: center;
justify-content: center;
padding: 20px;
z-index: 1000;
}
.modal-content {
background: #fff;
border-radius: 12px;
padding: 20px 24px;
max-width: 520px;
width: 100%;
box-shadow: 0 12px 50px rgba(0, 0, 0, 0.2);
position: relative;
font-size: 13px;
line-height: 1.4;
}
.modal-content h2 {
margin-top: 0;
font-size: 16px;
}
.modal-content p {
margin: 8px 0;
}
.modal-close {
position: absolute;
top: 8px;
right: 10px;
border: none;
background: transparent;
font-size: 20px;
cursor: pointer;
color: var(--muted);
}
.modal-close:hover {
color: var(--ink);
}
.data-cell {
white-space: normal;
position: relative;
}
.data-cell[data-env]:hover::after {
content: attr(data-env);
position: absolute;
top: 50%;
transform: translateY(-50%);
left: 50%;
transform: translate(-50%, -120%);
background: rgba(16, 24, 40, 0.92);
color: #fff;
padding: 4px 8px;
border-radius: 6px;
font-size: 11px;
white-space: nowrap;
pointer-events: none;
z-index: 5;
}
.data-cell[data-env]:hover::before {
content: "";
position: absolute;
top: 50%;
left: 50%;
transform: translate(-50%, -30%);
border: 6px solid transparent;
border-top-color: rgba(16, 24, 40, 0.92);
pointer-events: none;
z-index: 5;
}
.data-cell .measure,
.data-cell .std {
white-space: nowrap;
}
.row-actions {
display: flex;
gap: 6px;
margin-top: 4px;
flex-wrap: wrap;
}
.row-action-btn {
border: none;
background: transparent;
color: var(--accent);
font-size: 11px;
padding: 0;
cursor: pointer;
text-decoration: underline;
text-underline-offset: 2px;
}
.row-action-btn:hover {
color: #0d3fb8;
}
td.model .meta {
font-size: 10px;
color: var(--muted);
}
tbody tr:nth-child(even) td {
background: #fafbff;
}
.measure {
font-feature-settings: "tnum";
font-size: 12px;
font-weight: 600;
}
.std {
color: var(--muted);
font-size: 10px;
}
.winner-list {
display: flex;
flex-wrap: wrap;
gap: 2px;
}
.winner-pill {
display: inline-flex;
align-items: center;
padding: 2px 6px;
border-radius: 999px;
font-size: 10px;
background: #dbeafe;
color: #1e3a8a;
margin: 1px;
white-space: nowrap;
}
.cell-error {
color: var(--warn);
}
.cell-empty {
color: #c3c7d1;
}
.best {
background: var(--winner-bg) !important;
color: var(--winner-ink);
}
td.best .measure,
td.best .std {
color: var(--winner-ink);
}
.resize-handle {
position: absolute;
top: 0;
right: 0;
width: 6px;
height: 100%;
cursor: col-resize;
}
.resize-handle::after {
content: "";
position: absolute;
inset: 0;
background: transparent;
}
th.backend-header {
cursor: grab;
white-space: nowrap;
}
th.backend-header.dragging {
opacity: 0.5;
}
th.backend-header.drop-target {
outline: 2px dashed var(--accent);
}
.resize-line {
width: 2px;
background: var(--accent);
pointer-events: none;
}
.resize-overlay {
position: absolute;
top: 0;
bottom: 0;
left: 0;
right: 0;
pointer-events: none;
}
.resize-bar {
position: absolute;
top: 0;
bottom: 0;
width: 6px;
cursor: col-resize;
pointer-events: auto;
background: transparent;
}
.tag {
display: inline-flex;
align-items: center;
padding: 0 6px;
border-radius: 999px;
background: #f1f5ff;
color: #1d4ed8;
font-size: 11px;
}
.range-wrap {
position: relative;
height: 32px;
}
.range-wrap input[type="range"] {
position: absolute;
inset: 0;
width: 100%;
background: transparent;
-webkit-appearance: none;
appearance: none;
pointer-events: none;
}
.range-wrap input[type="range"]::-webkit-slider-thumb {
pointer-events: auto;
-webkit-appearance: none;
width: 18px;
height: 18px;
border-radius: 50%;
background: var(--accent);
border: 2px solid #fff;
box-shadow: 0 0 3px rgba(0, 0, 0, 0.3);
}
.range-wrap input[type="range"]::-moz-range-thumb {
pointer-events: auto;
width: 18px;
height: 18px;
border-radius: 50%;
background: var(--accent);
border: 2px solid #fff;
}
.range-track {
position: absolute;
top: 50%;
left: 0;
right: 0;
height: 6px;
border-radius: 999px;
background: #e3e7f1;
transform: translateY(-50%);
pointer-events: none;
}
.range-values {
font-size: 11px;
color: var(--muted);
margin-top: 4px;
}
.modal-content code {
font-family: "JetBrains Mono", "SFMono-Regular", Consolas, monospace;
background: #f6f8fc;
padding: 1px 4px;
border-radius: 4px;
font-size: 12px;
}
+767
View File
@@ -0,0 +1,767 @@
const DEFAULT_CTX = "default";
const K_SIGMA = 1.0;
const MIN_TOL = 0.25;
const MODEL_COL_WIDTH = 180;
const WINNER_COL_WIDTH = 120;
const state = {
contexts: [],
contextMap: new Map(),
envs: [],
backendOrder: [],
columnWidths: {},
filters: {
search: "",
quant: "",
context: DEFAULT_CTX,
backends: new Set(),
sizeLo: null,
sizeHi: null,
},
ui: {},
sizeStats: { min: Infinity, max: -Infinity },
draggingEnv: null,
};
document.addEventListener("DOMContentLoaded", async () => {
cacheUI();
setupModals();
try {
const res = await fetch("results.json");
const data = await res.json();
prepareData(data?.runs || []);
initializeControls();
renderTables();
} catch (err) {
console.error("Failed to load results.json", err);
state.ui.stats.textContent = "Failed to load results.json";
}
});
function cacheUI() {
state.ui = {
search: document.getElementById("filter-search"),
quant: document.getElementById("filter-quant"),
contextChips: document.getElementById("context-chips"),
backendList: document.getElementById("backend-list"),
backendAll: document.getElementById("backend-all"),
backendNone: document.getElementById("backend-none"),
sizeLo: document.getElementById("sizeLo"),
sizeHi: document.getElementById("sizeHi"),
sizeTrack: document.getElementById("sizeTrack"),
sizeLoVal: document.getElementById("sizeLoVal"),
sizeHiVal: document.getElementById("sizeHiVal"),
stats: document.getElementById("stats-line"),
resetBtn: document.getElementById("reset-layout"),
tables: document.getElementById("tables"),
hipblasModalOpen: document.getElementById("hipblas-modal-open"),
hipblasModal: document.getElementById("hipblas-modal"),
hipblasModalClose: document.getElementById("hipblas-modal-close"),
rpcModalOpen: document.getElementById("rpc-modal-open"),
rpcModal: document.getElementById("rpc-modal"),
rpcModalClose: document.getElementById("rpc-modal-close"),
rocwmmaModalOpen: document.getElementById("rocwmma-modal-open"),
rocwmmaModal: document.getElementById("rocwmma-modal"),
rocwmmaModalClose: document.getElementById("rocwmma-modal-close"),
rocwmmaImprModalOpen: document.getElementById("rocwmma-impr-modal-open"),
rocwmmaImprModal: document.getElementById("rocwmma-impr-modal"),
rocwmmaImprModalClose: document.getElementById("rocwmma-impr-modal-close"),
};
}
function setupModals() {
const modalConfigs = [
{
open: state.ui.hipblasModalOpen,
modal: state.ui.hipblasModal,
close: state.ui.hipblasModalClose,
},
{
open: state.ui.rpcModalOpen,
modal: state.ui.rpcModal,
close: state.ui.rpcModalClose,
},
{
open: state.ui.rocwmmaModalOpen,
modal: state.ui.rocwmmaModal,
close: state.ui.rocwmmaModalClose,
},
{
open: state.ui.rocwmmaImprModalOpen,
modal: state.ui.rocwmmaImprModal,
close: state.ui.rocwmmaImprModalClose,
},
];
modalConfigs.forEach(({ open, modal, close }) => {
if (!open || !modal) return;
const openModal = () => modal.classList.remove("hidden");
const closeModal = () => modal.classList.add("hidden");
open.addEventListener("click", openModal);
close?.addEventListener("click", closeModal);
modal.addEventListener("click", (e) => {
if (e.target === modal) closeModal();
});
document.addEventListener("keydown", (e) => {
if (e.key === "Escape" && !modal.classList.contains("hidden")) {
closeModal();
}
});
});
}
function prepareData(runs) {
const contextMap = new Map();
const envSet = new Set();
const quantSet = new Set();
for (const run of runs) {
const test = normalizeTest(run.test);
if (!test || !run.env) continue;
const contextKey = run.context || DEFAULT_CTX;
const env = run.env;
envSet.add(env);
if (run.quant) quantSet.add(run.quant.toUpperCase());
const ctx = ensureContext(contextMap, contextKey, run.context_tokens);
const testEntry = ensureTest(ctx, test.original);
const modelName = run.model_clean || run.model;
const row = ensureModel(testEntry, modelName, run);
row.backends[env] = {
mean: typeof run.tps_mean === "number" ? run.tps_mean : null,
std: typeof run.tps_std === "number" ? run.tps_std : null,
error: Boolean(run.error),
error_type: run.error_type || null,
};
}
state.contextMap = contextMap;
state.contexts = [...contextMap.values()].sort((a, b) => {
if (a.key === DEFAULT_CTX) return -1;
if (b.key === DEFAULT_CTX) return 1;
if (a.tokens && b.tokens) return a.tokens - b.tokens;
if (a.tokens) return -1;
if (b.tokens) return 1;
return a.key.localeCompare(b.key);
});
state.envs = [...envSet].sort();
state.backendOrder = [...state.envs];
state.columnWidths = Object.fromEntries(state.envs.map((env) => [env, 120]));
state.quantOptions = [...quantSet].sort();
state.filters.context = state.contexts[0]?.key || DEFAULT_CTX;
state.filters.backends = new Set(state.envs);
}
function ensureContext(map, key, tokens) {
if (!map.has(key)) {
map.set(key, {
key,
label: formatContextLabel(key, tokens),
tokens: tokens ?? null,
tests: new Map(),
});
} else if (tokens && !map.get(key).tokens) {
const ctx = map.get(key);
ctx.tokens = tokens;
ctx.label = formatContextLabel(key, tokens);
}
return map.get(key);
}
function ensureTest(ctx, testName) {
if (!ctx.tests.has(testName)) {
ctx.tests.set(testName, {
name: testName,
models: new Map(),
});
}
return ctx.tests.get(testName);
}
function ensureModel(testEntry, modelName, run) {
if (!testEntry.models.has(modelName)) {
testEntry.models.set(modelName, {
model: modelName,
quant: (run.quant || "Unknown").toUpperCase(),
sizeB: run.name_params_b ?? run.params_b ?? null,
backends: {},
isRpc: Boolean(run.rpc),
search_blob: [modelName, run.quant, run.env, run.test]
.filter(Boolean)
.map((s) => s.toString().toLowerCase())
.join(" "),
});
}
const row = testEntry.models.get(modelName);
const sizeCandidate = run.name_params_b ?? run.params_b;
if (row.sizeB == null && typeof sizeCandidate === "number") {
row.sizeB = sizeCandidate;
}
if (typeof row.sizeB === "number") {
state.sizeStats.min = Math.min(state.sizeStats.min, row.sizeB);
state.sizeStats.max = Math.max(state.sizeStats.max, row.sizeB);
}
if (run.rpc) {
row.isRpc = true;
if (!row.search_blob.includes("rpc")) {
row.search_blob = `${row.search_blob} rpc`;
}
}
return row;
}
function initializeControls() {
const { quant, contextChips, backendList, search, resetBtn, sizeLo, sizeHi } = state.ui;
quant.innerHTML = "";
const anyOpt = document.createElement("option");
anyOpt.value = "";
anyOpt.textContent = "Any";
quant.appendChild(anyOpt);
state.quantOptions.forEach((q) => {
const opt = document.createElement("option");
opt.value = q;
opt.textContent = q;
quant.appendChild(opt);
});
contextChips.innerHTML = "";
state.contexts.forEach((ctx) => {
const btn = document.createElement("button");
btn.type = "button";
btn.className = "chip" + (ctx.key === state.filters.context ? " active" : "");
btn.dataset.context = ctx.key;
btn.textContent = ctx.label;
contextChips.appendChild(btn);
});
renderBackendList();
setupSizeSlider();
search.addEventListener("input", (e) => {
state.filters.search = (e.target.value || "").trim().toLowerCase();
renderTables();
});
quant.addEventListener("change", (e) => {
state.filters.quant = e.target.value;
renderTables();
});
contextChips.addEventListener("click", (e) => {
const btn = e.target.closest("button[data-context]");
if (!btn) return;
state.filters.context = btn.dataset.context;
[...contextChips.querySelectorAll("button")].forEach((b) => b.classList.toggle("active", b === btn));
renderTables();
});
backendList.addEventListener("change", (e) => {
const checkbox = e.target.closest("input[data-env]");
if (!checkbox) return;
const env = checkbox.dataset.env;
if (checkbox.checked) {
state.filters.backends.add(env);
} else {
state.filters.backends.delete(env);
}
renderTables();
});
state.ui.backendAll.addEventListener("click", () => {
state.filters.backends = new Set(state.envs);
renderBackendList();
renderTables();
});
state.ui.backendNone.addEventListener("click", () => {
state.filters.backends = new Set();
renderBackendList();
renderTables();
});
sizeLo.addEventListener("input", () => updateSizeUI(true));
sizeHi.addEventListener("input", () => updateSizeUI(true));
resetBtn.addEventListener("click", () => {
state.filters.search = "";
state.filters.quant = "";
state.filters.context = state.contexts[0]?.key || DEFAULT_CTX;
state.filters.backends = new Set(state.envs);
search.value = "";
quant.value = "";
[...contextChips.querySelectorAll("button")].forEach((btn) =>
btn.classList.toggle("active", btn.dataset.context === state.filters.context)
);
renderBackendList();
setupSizeSlider();
renderTables();
});
}
function renderBackendList() {
const container = state.ui.backendList;
container.innerHTML = "";
state.backendOrder.forEach((env) => {
const label = document.createElement("label");
label.className = "backend-item";
const checkbox = document.createElement("input");
checkbox.type = "checkbox";
checkbox.dataset.env = env;
checkbox.checked = state.filters.backends.has(env);
label.appendChild(checkbox);
const baseSpan = document.createElement("span");
const { base, tags } = splitEnvName(env);
baseSpan.textContent = base;
label.appendChild(baseSpan);
tags.forEach((tag) => {
const pill = document.createElement("span");
pill.className = "tag";
pill.textContent = tag;
const safeTag = tag.replace(/[^a-z0-9]+/gi, "-").toLowerCase();
pill.classList.add(`tag-${safeTag}`);
label.appendChild(pill);
});
container.appendChild(label);
});
}
function setupSizeSlider() {
const { sizeLo, sizeHi } = state.ui;
const minRaw = state.sizeStats.min === Infinity ? 0 : Math.floor(state.sizeStats.min || 0);
const maxRaw = state.sizeStats.max === -Infinity ? 0 : Math.ceil(state.sizeStats.max || 0);
const minB = Math.max(0, minRaw);
const maxB = Math.max(minB, maxRaw);
[sizeLo, sizeHi].forEach((inp) => {
inp.min = minB;
inp.max = maxB;
inp.step = 1;
});
sizeLo.value = minB;
sizeHi.value = maxB;
sizeLo.style.zIndex = 2;
sizeHi.style.zIndex = 1;
updateSizeUI(false);
}
function updateSizeUI(triggerRender) {
const { sizeLo, sizeHi, sizeLoVal, sizeHiVal, sizeTrack } = state.ui;
if (+sizeLo.value > +sizeHi.value) {
if (document.activeElement === sizeLo) {
sizeHi.value = sizeLo.value;
} else {
sizeLo.value = sizeHi.value;
}
}
sizeLo.style.zIndex = +sizeLo.value >= +sizeHi.max - 1 ? 4 : 2;
sizeHi.style.zIndex = +sizeHi.value <= +sizeLo.min + 1 ? 3 : 1;
state.filters.sizeLo = +sizeLo.value;
state.filters.sizeHi = +sizeHi.value;
sizeLoVal.textContent = formatSizeLabel(state.filters.sizeLo);
sizeHiVal.textContent = formatSizeLabel(state.filters.sizeHi);
const range = (sizeHi.max - sizeLo.min) || 1;
const minB = +sizeLo.min;
const start = ((state.filters.sizeLo - minB) / range) * 100;
const end = ((state.filters.sizeHi - minB) / range) * 100;
sizeTrack.style.background = `linear-gradient(to right, #e3e7f1 ${start}%, var(--accent) ${start}%, var(--accent) ${end}%, #e3e7f1 ${end}%)`;
if (triggerRender) renderTables();
}
function renderTables() {
const ctx = state.contextMap.get(state.filters.context);
if (!ctx) {
state.ui.tables.innerHTML = "<p>No data for this context.</p>";
state.ui.stats.textContent = "0 rows";
return;
}
const backendList = state.backendOrder.filter((env) => state.filters.backends.has(env));
const tests = [...ctx.tests.values()].sort((a, b) => a.name.localeCompare(b.name));
const frag = document.createDocumentFragment();
let totalRows = 0;
for (const test of tests) {
const models = filterModels(test.models);
if (!models.length) continue;
totalRows += models.length;
const block = document.createElement("div");
block.className = "test-block";
const heading = document.createElement("h2");
heading.textContent = `${test.name.toUpperCase()} — tokens/second`;
block.appendChild(heading);
const tableWrap = document.createElement("div");
tableWrap.className = "table-wrap";
const scroller = document.createElement("div");
scroller.className = "table-scroll";
const modelsWithWinners = models.map((model) => {
const winners = computeWinners(model, backendList);
return { ...model, _cachedWinners: winners };
});
const table = buildSingleTable(modelsWithWinners, backendList);
scroller.appendChild(table);
tableWrap.appendChild(scroller);
block.appendChild(tableWrap);
setupResizeOverlay(scroller, backendList, table);
frag.appendChild(block);
}
state.ui.tables.innerHTML = "";
if (frag.childNodes.length) {
state.ui.tables.appendChild(frag);
} else {
state.ui.tables.innerHTML = "<p>No models match the current filters.</p>";
}
state.ui.stats.textContent = `Showing ${totalRows.toLocaleString()} model rows across ${backendList.length} backends`;
}
function buildSingleTable(models, backendList) {
const table = document.createElement("table");
const colgroup = document.createElement("colgroup");
const colModel = document.createElement("col");
colModel.style.width = `${MODEL_COL_WIDTH}px`;
colgroup.appendChild(colModel);
const colWinner = document.createElement("col");
colWinner.style.width = `${WINNER_COL_WIDTH}px`;
colgroup.appendChild(colWinner);
backendList.forEach((env) => {
const col = document.createElement("col");
col.style.width = `${state.columnWidths[env] || 120}px`;
col.dataset.env = env;
colgroup.appendChild(col);
});
table.appendChild(colgroup);
const thead = document.createElement("thead");
const headRow = document.createElement("tr");
headRow.appendChild(makeHeaderCell("Model", "model"));
headRow.appendChild(makeHeaderCell("Winner", "winner"));
backendList.forEach((env) => {
const th = makeHeaderCell(env, "backend-header");
attachHeaderInteractions(th, env);
headRow.appendChild(th);
});
thead.appendChild(headRow);
table.appendChild(thead);
const tbody = document.createElement("tbody");
models.forEach((model) => {
const tr = document.createElement("tr");
const tdModel = document.createElement("td");
tdModel.className = "model";
const head = document.createElement("div");
head.className = "model-head";
const nameSpan = document.createElement("span");
nameSpan.className = "model-name";
nameSpan.textContent = model.model;
head.appendChild(nameSpan);
if (model.isRpc) {
const pill = document.createElement("span");
pill.className = "model-pill model-pill-rpc";
pill.title = "Run executed via llama.cpp RPC across two servers";
pill.textContent = "RPC · dual server";
head.appendChild(pill);
}
tdModel.appendChild(head);
const meta = document.createElement("div");
meta.className = "meta";
meta.textContent = `${model.quant} · ${formatSize(model.sizeB)}`;
tdModel.appendChild(meta);
const actionWrap = document.createElement("div");
actionWrap.className = "row-actions";
const btnDesc = document.createElement("button");
btnDesc.type = "button";
btnDesc.className = "row-action-btn";
btnDesc.textContent = "Sort ↓";
btnDesc.addEventListener("click", (e) => {
e.preventDefault();
sortBackendsByModel(model, "desc");
});
const btnAsc = document.createElement("button");
btnAsc.type = "button";
btnAsc.className = "row-action-btn";
btnAsc.textContent = "Sort ↑";
btnAsc.addEventListener("click", (e) => {
e.preventDefault();
sortBackendsByModel(model, "asc");
});
actionWrap.appendChild(btnDesc);
actionWrap.appendChild(btnAsc);
tdModel.appendChild(actionWrap);
tr.appendChild(tdModel);
const tdWinner = document.createElement("td");
tdWinner.className = "winner";
if (model._cachedWinners.length) {
const wrap = document.createElement("div");
wrap.className = "winner-list";
wrap.innerHTML = model._cachedWinners.map((w) => `<span class="winner-pill">${w}</span>`).join("");
tdWinner.appendChild(wrap);
} else {
tdWinner.innerHTML = `<span class="cell-empty">—</span>`;
}
tr.appendChild(tdWinner);
backendList.forEach((env) => {
const td = document.createElement("td");
td.className = "data-cell";
td.dataset.env = env;
const cell = model.backends[env];
if (!cell) {
td.innerHTML = `<span class="cell-empty">—</span>`;
} else if (cell.error || cell.mean == null) {
td.innerHTML = `<span class="cell-error">⚠ ${cell.error_type || "error"}</span>`;
} else {
const isBest = model._cachedWinners.includes(env);
if (isBest) td.classList.add("best");
td.innerHTML = `<div class="measure">${cell.mean.toFixed(2)}</div><div class="std">± ${cell.std?.toFixed(2) ?? "—"}</div>`;
}
tr.appendChild(td);
});
tbody.appendChild(tr);
});
table.appendChild(tbody);
return table;
}
function makeHeaderCell(label, extra = "") {
const th = document.createElement("th");
th.textContent = label;
if (extra) th.className = extra;
return th;
}
function attachHeaderInteractions(th, env) {
const width = state.columnWidths[env] || 120;
th.style.width = `${width}px`;
th.style.minWidth = `${width}px`;
th.draggable = true;
th.addEventListener("dragstart", (e) => {
state.draggingEnv = env;
th.classList.add("dragging");
e.dataTransfer.effectAllowed = "move";
});
th.addEventListener("dragend", () => {
state.draggingEnv = null;
th.classList.remove("dragging");
document.querySelectorAll("th.backend-header.drop-target").forEach((el) => el.classList.remove("drop-target"));
});
th.addEventListener("dragover", (e) => {
if (!state.draggingEnv || state.draggingEnv === env) return;
e.preventDefault();
th.classList.add("drop-target");
});
th.addEventListener("dragleave", () => th.classList.remove("drop-target"));
th.addEventListener("drop", (e) => {
if (!state.draggingEnv || state.draggingEnv === env) return;
e.preventDefault();
moveBackend(state.draggingEnv, env);
th.classList.remove("drop-target");
});
const handle = document.createElement("span");
handle.className = "resize-handle";
handle.addEventListener("mousedown", (e) => startResize(e, env));
th.appendChild(handle);
}
function moveBackend(from, to) {
const order = state.backendOrder;
const fromIdx = order.indexOf(from);
const toIdx = order.indexOf(to);
if (fromIdx === -1 || toIdx === -1) return;
const [col] = order.splice(fromIdx, 1);
order.splice(toIdx, 0, col);
renderBackendList();
renderTables();
}
function filterModels(modelsMap) {
const models = [];
for (const model of modelsMap.values()) {
if (state.filters.search && !model.search_blob.includes(state.filters.search)) continue;
if (state.filters.quant && model.quant !== state.filters.quant) continue;
if (model.sizeB != null) {
if (state.filters.sizeLo != null && model.sizeB < state.filters.sizeLo - 1e-6) continue;
if (state.filters.sizeHi != null && model.sizeB > state.filters.sizeHi + 1e-6) continue;
}
models.push(model);
}
models.sort((a, b) => a.model.localeCompare(b.model));
return models;
}
function computeWinners(model, backends) {
const values = [];
backends.forEach((env) => {
const entry = model.backends[env];
if (entry && !entry.error && typeof entry.mean === "number") {
values.push({
env,
mean: entry.mean,
std: typeof entry.std === "number" ? entry.std : 0,
});
}
});
if (!values.length) return [];
let best = values[0];
for (const v of values) if (v.mean > best.mean) best = v;
const winners = [];
for (const v of values) {
const pooled = Math.sqrt((best.std || 0) ** 2 + (v.std || 0) ** 2);
const tol = Math.max(MIN_TOL, K_SIGMA * pooled);
if ((best.mean - v.mean) <= tol) winners.push(v.env);
}
return winners;
}
function normalizeTest(name) {
if (!name) return null;
return { key: name.toLowerCase(), original: name };
}
function formatContextLabel(key, tokens) {
if (key === DEFAULT_CTX) return "Default window";
if (tokens) return `ctx ${tokens.toLocaleString()}`;
return key;
}
function formatSize(size) {
if (size == null) return "—";
return `${Number(size).toFixed(1)}B`;
}
function formatSizeLabel(size) {
if (size >= 1000) return `${(size / 1000).toFixed(1)}kB`;
return `${Math.round(size)}B`;
}
function sortBackendsByModel(model, direction) {
const dir = direction === "asc" ? 1 : -1;
const order = [...state.backendOrder].sort((a, b) => {
const va = backendValue(model.backends[a], direction);
const vb = backendValue(model.backends[b], direction);
if (va === vb) return a.localeCompare(b);
return (va - vb) * dir;
});
state.backendOrder = order;
renderBackendList();
renderTables();
}
function backendValue(entry, direction) {
if (!entry || entry.error || typeof entry.mean !== "number") {
return direction === "asc" ? Number.POSITIVE_INFINITY : Number.NEGATIVE_INFINITY;
}
return entry.mean;
}
function splitEnvName(env) {
const canonical = env.replace(/_/g, ".");
const tagRegex = /-(rocwmma-improved|rocwmma|improved|hblt0)/gi;
const tags = [];
let match;
while ((match = tagRegex.exec(canonical)) !== null) {
tags.push(match[1].toLowerCase());
}
const base = canonical.replace(tagRegex, "");
return { base, tags };
}
function startResize(event, env) {
event.preventDefault();
event.stopPropagation();
const column = state.columnWidths[env] || 120;
const startX = event.clientX;
const shellRect = state.ui.tables.getBoundingClientRect();
const guide = document.createElement("div");
guide.className = "resize-line";
guide.style.position = "fixed";
guide.style.top = `${shellRect.top}px`;
guide.style.bottom = `${window.innerHeight - shellRect.bottom}px`;
guide.style.left = `${startX}px`;
guide.style.width = "2px";
guide.style.background = "var(--accent)";
guide.style.zIndex = "10";
document.body.appendChild(guide);
let nextWidth = column;
const onMove = (e) => {
const delta = e.clientX - startX;
nextWidth = Math.max(80, column + delta);
guide.style.left = `${e.clientX}px`;
};
const onUp = () => {
document.removeEventListener("mousemove", onMove);
document.removeEventListener("mouseup", onUp);
guide.remove();
state.columnWidths[env] = nextWidth;
renderTables();
};
document.addEventListener("mousemove", onMove);
document.addEventListener("mouseup", onUp);
}
function setupResizeOverlay(tableWrap, backendList, table) {
let overlay = tableWrap.querySelector(".resize-overlay");
if (!overlay) {
overlay = document.createElement("div");
overlay.className = "resize-overlay";
tableWrap.appendChild(overlay);
} else {
overlay.innerHTML = "";
}
overlay.style.width = `${tableWrap.clientWidth}px`;
overlay.style.height = `${table.offsetHeight}px`;
const bars = [];
let offset = MODEL_COL_WIDTH + WINNER_COL_WIDTH;
backendList.forEach((env) => {
const width = state.columnWidths[env] || 120;
const bar = document.createElement("div");
bar.className = "resize-bar";
bar.dataset.env = env;
bar.addEventListener("mousedown", (e) => startResize(e, env));
overlay.appendChild(bar);
bars.push({ bar, offset, width, env });
offset += width;
});
const positionBars = () => {
bars.forEach(({ bar, offset, width }) => {
const left = offset + width - 3 - tableWrap.scrollLeft;
bar.style.left = `${left}px`;
});
};
positionBars();
if (tableWrap._overlayScroll) {
tableWrap.removeEventListener("scroll", tableWrap._overlayScroll);
}
const onScroll = () => positionBars();
tableWrap.addEventListener("scroll", onScroll);
tableWrap._overlayScroll = onScroll;
if (tableWrap._overlayResize) {
tableWrap._overlayResize.disconnect();
}
const resizeObserver = new ResizeObserver(() => {
overlay.style.width = `${tableWrap.clientWidth}px`;
overlay.style.height = `${table.offsetHeight}px`;
positionBars();
});
resizeObserver.observe(tableWrap);
tableWrap._overlayResize = resizeObserver;
}
+164
View File
@@ -0,0 +1,164 @@
# AMD Strix Halo — llama.cpp Toolboxes (Benchmarks)
**Interactive results:** https://kyuz0.github.io/amd-strix-halo-toolboxes/
## Table of Contents
- [Benchmark methodology](#benchmark-methodology)
- [Summary of current dataset (Flash Attention ON)](#summary-of-current-dataset-flash-attention-on)
- [Placement counts](#placement-counts)
- [Pairwise head-to-head wins](#pairwise-head-to-head-wins)
- [Average ranks](#average-ranks)
- [Analyses by feature](#analyses-by-feature)
- [Impact of Flash Attention](#impact-of-flash-attention)
- [Impact of ROCWMMA](#impact-of-rocwmma)
- [Impact of hipBLASLt](#impact-of-hipblaslt)
- [Vulkan: AMDVLK vs RADV](#vulkan-amdvlk-vs-radv)
- [Recommendations](#recommendations)
- [Winner calculation](#winner-calculation)
---
## Benchmark methodology
- **pp512** — prompt processing throughput (tokens/sec, prefill)
- **tg128** — token generation throughput (tokens/sec, interactive)
- Each backend tested twice per model: `-fa 0` and `-fa 1`
- Winners per model/test are **margin-aware**; multiple winners are possible when mean±σ overlap
- Built from the same llama.cpp commit for consistency
**Backends in this dataset:** ROCm 7 RC + ROCWMMA + hipBLASLt, ROCm 7 RC (hipBLASLt), ROCm 7 RC (hipBLASLt OFF), ROCm 7 RC + ROCWMMA (hipBLASLt OFF), ROCm 6.4.4 (hipBLASLt), ROCm 6.4.4 (hipBLASLt OFF), ROCm 6.4.4 + ROCWMMA (hipBLASLt), ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF), Vulkan AMDVLK, Vulkan RADV
**ROCm 7 hipBLASLt policy:** Toolboxes ship with **hipBLASLt enabled** by default (`ROCBLAS_USE_HIPBLASLT=1`). The benchmark script also runs **hipBLASLt OFF** variants (`-hblt0`) to measure its effect.
---
## Summary of current dataset (Flash Attention ON)
### Placement counts
**Prompt Processing (pp512)**
| Backend | 1st | 2nd | 3rd |
| --- | ---: | ---: | ---: |
| ROCm 6.4.4 (hipBLASLt) | 6 | 2 | 2 |
| Vulkan AMDVLK | 6 | 1 | 0 |
| ROCm 6.4.4 (hipBLASLt OFF) | 3 | 2 | 3 |
| Vulkan RADV | 1 | 2 | 0 |
| ROCm 7 RC (hipBLASLt) | 1 | 1 | 1 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 0 | 5 | 4 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt) | 0 | 4 | 2 |
| ROCm 7 RC (hipBLASLt OFF) | 0 | 0 | 2 |
| ROCm 7 RC + ROCWMMA + hipBLASLt | 0 | 0 | 3 |
**Token Generation (tg128)**
| Backend | 1st | 2nd | 3rd |
| --- | ---: | ---: | ---: |
| Vulkan RADV | 10 | 1 | 2 |
| Vulkan AMDVLK | 3 | 10 | 0 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 2 | 3 | 7 |
| ROCm 6.4.4 (hipBLASLt) | 1 | 4 | 3 |
| ROCm 6.4.4 (hipBLASLt OFF) | 1 | 3 | 5 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt) | 1 | 2 | 6 |
| ROCm 7 RC (hipBLASLt) | 1 | 0 | 1 |
| ROCm 7 RC (hipBLASLt OFF) | 0 | 1 | 1 |
| ROCm 7 RC + ROCWMMA + hipBLASLt | 0 | 1 | 1 |
| ROCm 7 RC + ROCWMMA (hipBLASLt OFF) | 0 | 1 | 1 |
### Pairwise head-to-head wins
For any model+quant where both backends succeeded, this counts who was faster (ties when equal).
| Comparison | Test | A wins | B wins | Ties | Total |
| --- | --- | ---: | ---: | ---: | ---: |
| ROCm 7 RC + ROCWMMA + hipBLASLt vs Vulkan AMDVLK | pp512 | 9 | 7 | 0 | 16 |
| ROCm 7 RC + ROCWMMA + hipBLASLt vs Vulkan AMDVLK | tg128 | 2 | 14 | 0 | 16 |
| ROCm 7 RC + ROCWMMA + hipBLASLt vs Vulkan RADV | pp512 | 14 | 3 | 0 | 17 |
| ROCm 7 RC + ROCWMMA + hipBLASLt vs Vulkan RADV | tg128 | 4 | 12 | 1 | 17 |
| Vulkan AMDVLK vs Vulkan RADV | pp512 | 12 | 4 | 0 | 16 |
| Vulkan AMDVLK vs Vulkan RADV | tg128 | 5 | 11 | 0 | 16 |
### Average ranks
**Prompt Processing (pp512)**
| Backend | Avg Rank (↓ is better) |
| --- | ---: |
| Vulkan AMDVLK | 1.14 |
| ROCm 6.4.4 (hipBLASLt) | 1.6 |
| Vulkan RADV | 1.67 |
| ROCm 6.4.4 (hipBLASLt OFF) | 2.0 |
| ROCm 7 RC (hipBLASLt) | 2.0 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt) | 2.33 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 2.44 |
| ROCm 7 RC (hipBLASLt OFF) | 3.0 |
| ROCm 7 RC + ROCWMMA + hipBLASLt | 3.0 |
**Token Generation (tg128)**
| Backend | Avg Rank (↓ is better) |
| --- | ---: |
| Vulkan RADV | 1.38 |
| Vulkan AMDVLK | 1.77 |
| ROCm 7 RC (hipBLASLt) | 2.0 |
| ROCm 6.4.4 (hipBLASLt) | 2.25 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 2.42 |
| ROCm 6.4.4 (hipBLASLt OFF) | 2.44 |
| ROCm 7 RC + ROCWMMA + hipBLASLt | 2.5 |
| ROCm 7 RC (hipBLASLt OFF) | 2.5 |
| ROCm 7 RC + ROCWMMA (hipBLASLt OFF) | 2.5 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt) | 2.56 |
---
## Analyses by feature
### Impact of Flash Attention
Median % change when **Flash Attention ON vs OFF**, paired by model+quant, per backend:
| Backend | pp512 Δ% (median, min..max, n) | tg128 Δ% (median, min..max, n) |
| --- | --- | --- |
| ROCm 7 RC + ROCWMMA + hipBLASLt | 11.4% (4.2..34.1), n=17 | -0.5% (-8.8..0.8), n=17 |
| ROCm 7 RC (hipBLASLt) | 11.7% (-23.0..25.6), n=14 | -1.1% (-8.7..1.0), n=14 |
| ROCm 7 RC (hipBLASLt OFF) | 6.8% (2.1..18.4), n=15 | -0.8% (-9.0..0.5), n=15 |
| ROCm 7 RC + ROCWMMA (hipBLASLt OFF) | 6.3% (-5.5..17.4), n=16 | -0.8% (-15.1..0.6), n=16 |
| ROCm 6.4.4 (hipBLASLt) | 8.3% (5.6..20.8), n=17 | 0.8% (-3.0..2.6), n=17 |
| ROCm 6.4.4 (hipBLASLt OFF) | 7.2% (-0.5..19.5), n=17 | 1.1% (-2.9..2.7), n=17 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt) | 7.1% (5.0..19.9), n=17 | 0.9% (-2.8..2.8), n=17 |
| ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 6.5% (2.7..18.6), n=17 | 1.1% (-2.7..3.4), n=17 |
| Vulkan AMDVLK | 1.3% (-10.8..27.8), n=16 | -1.2% (-6.8..0.1), n=16 |
| Vulkan RADV | 4.8% (-0.5..20.1), n=17 | -0.1% (-2.1..2.0), n=17 |
### Impact of ROCWMMA
| Context | Test | Compared Envs | Pairs | Median Δ% |
| --- | --- | --- | ---: | ---: |
| ROCm 7 RC (hipBLASLt) | pp512 | ROCm 7 RC + ROCWMMA + hipBLASLt vs ROCm 7 RC (hipBLASLt) | 15 | -0.0% |
| ROCm 7 RC (hipBLASLt) | tg128 | ROCm 7 RC + ROCWMMA + hipBLASLt vs ROCm 7 RC (hipBLASLt) | 15 | 0.0% |
| ROCm 7 RC (hipBLASLt OFF) | pp512 | ROCm 7 RC + ROCWMMA (hipBLASLt OFF) vs ROCm 7 RC (hipBLASLt OFF) | 17 | -0.2% |
| ROCm 7 RC (hipBLASLt OFF) | tg128 | ROCm 7 RC + ROCWMMA (hipBLASLt OFF) vs ROCm 7 RC (hipBLASLt OFF) | 17 | 0.0% |
| ROCm 6.4.4 (hipBLASLt) | pp512 | ROCm 6.4.4 + ROCWMMA (hipBLASLt) vs ROCm 6.4.4 (hipBLASLt) | 17 | -0.4% |
| ROCm 6.4.4 (hipBLASLt) | tg128 | ROCm 6.4.4 + ROCWMMA (hipBLASLt) vs ROCm 6.4.4 (hipBLASLt) | 17 | 0.0% |
| ROCm 6.4.4 (hipBLASLt OFF) | pp512 | ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) vs ROCm 6.4.4 (hipBLASLt OFF) | 17 | -0.5% |
| ROCm 6.4.4 (hipBLASLt OFF) | tg128 | ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) vs ROCm 6.4.4 (hipBLASLt OFF) | 17 | -0.1% |
### Impact of hipBLASLt
| Context | Test | Compared Envs | Pairs | Median Δ% |
| --- | --- | --- | ---: | ---: |
| ROCm 7 RC (no ROCWMMA) | pp512 | ROCm 7 RC (hipBLASLt) vs ROCm 7 RC (hipBLASLt OFF) | 15 | -0.2% |
| ROCm 7 RC (no ROCWMMA) | tg128 | ROCm 7 RC (hipBLASLt) vs ROCm 7 RC (hipBLASLt OFF) | 15 | 0.0% |
| ROCm 7 RC + ROCWMMA | pp512 | ROCm 7 RC + ROCWMMA + hipBLASLt vs ROCm 7 RC + ROCWMMA (hipBLASLt OFF) | 17 | -0.1% |
| ROCm 7 RC + ROCWMMA | tg128 | ROCm 7 RC + ROCWMMA + hipBLASLt vs ROCm 7 RC + ROCWMMA (hipBLASLt OFF) | 17 | 0.0% |
| ROCm 6.4.4 (no ROCWMMA) | pp512 | ROCm 6.4.4 (hipBLASLt) vs ROCm 6.4.4 (hipBLASLt OFF) | 17 | 0.0% |
| ROCm 6.4.4 (no ROCWMMA) | tg128 | ROCm 6.4.4 (hipBLASLt) vs ROCm 6.4.4 (hipBLASLt OFF) | 17 | 0.0% |
| ROCm 6.4.4 + ROCWMMA | pp512 | ROCm 6.4.4 + ROCWMMA (hipBLASLt) vs ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 17 | -0.3% |
| ROCm 6.4.4 + ROCWMMA | tg128 | ROCm 6.4.4 + ROCWMMA (hipBLASLt) vs ROCm 6.4.4 + ROCWMMA (hipBLASLt OFF) | 17 | 0.0% |
### Vulkan: AMDVLK vs RADV
Head-to-head wins with selected Flash Attention filter:
| Test | AMDVLK wins | RADV wins | Ties | Total |
| --- | ---: | ---: | ---: | ---: |
| pp512 | 12 | 4 | 0 | 16 |
| tg128 | 5 | 11 | 0 | 16 |
---
## Recommendations
- **Fastest prompt processing:** Vulkan AMDVLK, ROCm 6.4.4 (hipBLASLt) (most 1st-place finishes with selected Flash Attention filter).
- **Fastest token generation:** Vulkan RADV (most 1st-place finishes with selected Flash Attention filter).
- **Balanced choice:** Vulkan AMDVLK (consistently near the top across PP/TG).
---
## Winner calculation
A backend is counted as a winner if its mean throughput is within the best backend’s pooled ± error margin for that model/test type. This treats results within measurement noise as ties instead of false losses.
+75
View File
@@ -0,0 +1,75 @@
# Building Containers Locally
If you want to build or customize the toolbox containers yourself (rather than using the pre-built Docker Hub images), this guide explains the process. Local builds are useful if you want to:
* Use a patched or forked version of llama.cpp
* Add additional tools or libraries
* Change the Fedora base image (Rawhide vs. stable)
* Audit every installed dependency
---
## 1. Prerequisites
* **Podman** (recommended on Fedora) or **Docker** (also fine)
---
## 2. Build an Image
Each backend has its own subdirectory and Dockerfile in `toolboxes/`.
**Example: Build the Vulkan RADV toolbox image**
```sh
cd toolboxes
podman build --no-cache -t llama-vulkan-radv -f Dockerfile.vulkan-radv .
```
**Example: Build the ROCm 6.4.2 toolbox image**
```sh
cd toolboxes
podman build --no-cache -t llama-rocm-6.4.2 -f Dockerfile.rocm-6.4.2 .
```
> You can use `docker build` if you prefer Docker.
---
## 3. Customizing the Build
* **llama.cpp version**: Change the `git clone` or `git checkout` line in the Dockerfile.
* **Extra dependencies**: Add them to the Dockerfile as needed.
* **Other customizations**: Install tools, patch scripts, or swap to a different base image.
---
## 4. Using the Custom Image with Toolbx
Create a new toolbox using your freshly built image:
```sh
toolbox create llama-vulkan-radv --image localhost/llama-vulkan-radv \
-- --device /dev/dri --group-add video --security-opt seccomp=unconfined
```
Replace the backend/image name and device/group options as needed (see main README Section 2.1).
---
## 5. Troubleshooting
* **Build fails (ROCm images especially):** Try building with more memory or swap.
* **Toolbox can't access GPU:** Make sure you pass the correct device/group options.
---
## 6. References
* [Fedora Toolbox Documentation](https://docs.fedoraproject.org/en-US/fedora-silverblue/toolbox/)
* [Podman Build Reference](https://docs.podman.io/en/latest/markdown/podman-build.1.html)
* [Docker Build Reference](https://docs.docker.com/engine/reference/commandline/build/)
+119
View File
@@ -0,0 +1,119 @@
## How to use docker-compose instead of toolbox
## Table of Contents
1. [Vulkan AMDVLK](#1-vulkanamdvlk)
2. [ROCm-6.4.4+ROCWMMA](#2-rocm-644-rocwmma)
## 1. Vulkan(AMDVLK)
1. Select applicable backend Dockerfile from repo. Example:
https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/main/toolboxes/Dockerfile.vulkan-amdvlk
2. In the build file, change shell command to:
```
# shell
CMD ["/bin/bash", "-c", "llama-server --host $HOST --port $PORT -c $CONTEXT_LENGTH --temp $TEMPERATURE --jinja --no-mmap -ngl $NGL -fa $FA -m $MODEL_PATH"]
```
3. Build container with:
```
docker build -f Dockerfile.vulkan-amdvlk -t vulkan-amdvlk:1.0 .
```
4. Download your model files to a directory. We will mount this from the container. I use:
```
/mnt/models
```
5. Create your docker compose, using this template. Change the ports and paths as needed.
```
services:
gpt-oss-120b:
container_name: gpt-oss-120b
image: vulkan-amdvlk:1.0
ports:
- "8069:8069"
volumes:
- /mnt/models:/mnt/models
devices:
- "/dev/dri:/dev/dri"
privileged: true
restart: unless-stopped
environment:
- HOST=0.0.0.0
- PORT=8069
- CONTEXT_LENGTH=120000
- TEMPERATURE=0.0
- MODEL_PATH=/mnt/models/gpt-oss-120b-UD-Q4_K_XL/gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf
- NGL=999
- FA=on
```
6. Start as usual.
```
docker compose up -d
```
## 2. ROCm-6.4.4-ROCWMMA
1. Select applicable backend Dockerfile from repo. Example:
https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/main/toolboxes/Dockerfile.rocm-6.4.4-rocwmma
3. In the build file, change shell command to:
```
# shell
CMD ["/bin/bash", "-c", "llama-server --host $HOST --port $PORT -c $CONTEXT_LENGTH --temp $TEMPERATURE --jinja --no-mmap -ngl $NGL -fa $FA -m $MODEL_PATH"]
```
3. Build container with:
```
docker build -f Dockerfile.rocm-6.4.4-rocwmma -t rocm-6.4.4-rocwmma:1.0 .
```
4. Download your model files to a directory. We will mount this from the container. I use:
```
/mnt/models
```
5. Create your docker compose, using this template. Change the ports and paths as needed.
```
services:
gpt-oss-120b:
container_name: gpt-oss-120b
image: rocm-6.4.4-rocwmma:1.0
ports:
- "8069:8069"
volumes:
- /mnt/models:/mnt/models
devices:
- "/dev/dri:/dev/dri"
- "/dev/kfd:/dev/kfd"
privileged: true
restart: unless-stopped
environment:
- HOST=0.0.0.0
- PORT=8069
- CONTEXT_LENGTH=120000
- TEMPERATURE=0.0
- MODEL_PATH=/mnt/models/gpt-oss-120b-UD-Q4_K_XL/gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf
- NGL=999
- FA=on
```
6. Start as usual.
```
docker compose up -d
```
+139
View File
@@ -0,0 +1,139 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>AMD Radeon AI PRO R9700 — Backend Benchmarks (Grid View)</title>
<link rel="stylesheet" href="assets/index2.css">
</head>
<body>
<header>
<h1>AMD Radeon AI PRO R9700 — Benchmark Grid</h1>
<p>AMD Radeon AI PRO R9700 · 32GB vRAM</p>
<p>Fedora 43 · Linux 6.17.8-300.fc43.x86_64 · llama.cpp build 1c398dc9e (7034)</p>
<p>Benchmarks captured 14 Nov 2025 · Repo: <a href="https://github.com/kyuz0/amd-r9700-toolboxes"
target="_blank" rel="noreferrer">kyuz0/amd-r9700-toolboxes</a></p>
<div class="legend">
<label>Legend</label>
<div class="legend-pills">
<button id="hipblas-modal-open" type="button" class="chip small legend-pill legend-pill-default">
hipBLASLt vs hblt0
</button>
<button id="rpc-modal-open" type="button" class="chip small legend-pill legend-pill-rpc">
RPC · dual server
</button>
<button id="rocwmma-modal-open" type="button" class="chip small legend-pill legend-pill-rocwmma">
rocWMMA
</button>
</div>
</div>
</header>
<section class="controls">
<div class="control">
<label for="filter-search">Search models</label>
<input id="filter-search" type="text" placeholder="e.g. llama, qwen, 30B…">
</div>
<div class="control">
<label for="filter-quant">Quant</label>
<select id="filter-quant">
<option value="">Any</option>
</select>
</div>
<div class="control grow slider-block">
<label>Context windows</label>
<div id="context-chips" class="chip-row tight"></div>
</div>
<div class="control grow slider-block">
<label>Model params (B)</label>
<div class="range-wrap">
<input type="range" id="sizeLo" step="1">
<input type="range" id="sizeHi" step="1">
<div class="range-track" id="sizeTrack"></div>
</div>
<div class="range-values">
<span id="sizeLoVal">0B</span> – <span id="sizeHiVal">0B</span>
</div>
</div>
</section>
<section class="panel compact">
<div class="panel-split">
<div class="backend-header">
<div class="backend-label">
<label>Backends</label>
<div class="backend-actions">
<button type="button" id="backend-all" class="chip small">All</button>
<button type="button" id="backend-none" class="chip small">None</button>
</div>
</div>
<div id="backend-list" class="backend-list"></div>
</div>
<div class="stats-box">
<div class="stat-line" id="stats-line">Loading…</div>
<button id="reset-layout" type="button" class="chip small">Reset filters</button>
</div>
</div>
</section>
<section class="panel compact" id="tables-panel">
<div id="tables"></div>
</section>
<div id="hipblas-modal" class="modal hidden" role="dialog" aria-modal="true" aria-labelledby="hipblas-title">
<div class="modal-content">
<button id="hipblas-modal-close" class="modal-close" aria-label="Close dialog">×</button>
<h2 id="hipblas-title">hipBLASLt &amp; hblt0 explained</h2>
<p>The ROCm toolboxes ship with <code>ROCBLAS_USE_HIPBLASLT=1</code> by default. This forces rocBLAS to
prefer
the hipBLASLt kernel library, which historically delivered the best throughput on gfx1201 (R9700).</p>
<p>Rows tagged with <code>__hblt0</code> were re-run with <code>ROCBLAS_USE_HIPBLASLT=0</code>, letting
rocBLAS
auto-select between hipBLASLt, Tensile, or other kernel providers. These runs show how performance
shifts when
the tuned hipBLASLt path is disabled.</p>
<p>hipBLASLt is AMD's LT (low-level tuned) matmul backend, optimized for transformer workloads. Disabling it
can
expose regressions or improvements depending on driver versions, so both configurations are published
for
comparison.</p>
</div>
</div>
<div id="rpc-modal" class="modal hidden" role="dialog" aria-modal="true" aria-labelledby="rpc-title">
<div class="modal-content">
<button id="rpc-modal-close" class="modal-close" aria-label="Close dialog">×</button>
<h2 id="rpc-title">RPC · dual server</h2>
<p>These results were produced with two R9700 systems (each 32&nbsp;GB)
connected over 5&nbsp;Gbps Ethernet. One runs <code>rpc-server</code> from llama.cpp; the other runs
<code>llama-bench --rpc</code>.
</p>
<p>This setup allows distributed inference, splitting large GGUF models across both machines. The metric
shows what
you can expect when latency is limited by the network and the workload is balanced between two RPC
participants.</p>
</div>
</div>
<div id="rocwmma-modal" class="modal hidden" role="dialog" aria-modal="true" aria-labelledby="rocwmma-title">
<div class="modal-content">
<button id="rocwmma-modal-close" class="modal-close" aria-label="Close dialog">×</button>
<h2 id="rocwmma-title">rocWMMA variants</h2>
<p>Backends labeled <code>-rocwmma</code> are rebuilt with AMD's rocWMMA library, which unlocks matrix
multiply
pipelines accelerated via wave matrix multiply-accumulate (WMMA) instructions.</p>
<p>rocWMMA kernels can significantly accelerate BF16/F16 workloads on RDNA3 but may trade stability or
memory
usage; comparing plain toolboxes against <code>-rocwmma</code> ones highlights the benefit or cost.</p>
</div>
</div>
<script src="assets/index2.js" type="module"></script>
</body>
</html>
+45422
View File
File diff suppressed because it is too large Load Diff
+89
View File
@@ -0,0 +1,89 @@
---
## docs/vram-estimator.md
---
# 1. Memory Planning with `gguf-vram-estimator.py`
Estimating memory requirements is critical when running large models on Strix Halo (or any GPU with limited RAM). It's not enough to check just the model file size: context length and runtime overheads matter.
This repo provides a tool, **`gguf-vram-estimator.py`**, which reads a `.gguf` model and prints the estimated VRAM needed for different context sizes.
**Why?**
* Helps decide what fits on 32GB, 64GB, 128GB, etc—especially with multi-shard models or large quantized files.
---
## 2. Usage
Make sure you have the estimator script (in `tools/`):
```sh
gguf-vram-estimator.py <path-to-model.gguf>
```
* Supply one or more context lengths to get the corresponding VRAM footprint.
* Handles multi-shard and single-shard models.
---
## 3. Examples
### 3.1 Llama-4-Scout 17B Q4\_K\_XL, up to 1M tokens
```
$ gguf-vram-estimator.py models/llama-4-scout-17b-16e/Q4_K_XL/Llama-4-Scout-17B-16E-Instruct-UD-Q4_K_XL-00001-of-00002.gguf --contexts 4096 32768 1048576
--- Model 'Llama-4-Scout-17B-16E-Instruct' ---
Max Context: 10,485,760 tokens
Model Size: 57.74 GiB
Incl. Overhead: 2.00 GiB
--- Memory Footprint Estimation ---
Context Size | Context Memory | Est. Total VRAM
---------------------------------------------------
4,096 | 1.88 GiB | 61.62 GiB
32,768 | 15.06 GiB | 74.80 GiB
1,048,576 | 49.12 GiB | 108.87 GiB
```
* **Takeaway:**
* Q4\_K quantization allows for a huge context in 128GB, but *processing 1M tokens will be extremely slow* (see benchmark: 200 tokens/sec prompt processing ⇒ almost 1.5 hours for a full 1M context fill).
---
### 3.2 Qwen3-235B Q3\_K XL, high context
```
$ gguf-vram-estimator.py models/qwen3-235B-Q3_K-XL/UD-Q3_K_XL/Qwen3-235B-A22B-Instruct-2507-UD-Q3_K_XL-00001-of-00003.gguf --contexts 65536 131072 262144
--- Memory Footprint Estimation ---
Context Size | Context Memory | Est. Total VRAM
---------------------------------------------------
65,536 | 11.75 GiB | 110.75 GiB
131,072 | 23.50 GiB | 122.50 GiB
262,144 | 47.00 GiB | 146.00 GiB
```
* **Takeaway:**
* With 128GB, you can go up to \~130k context on this Qwen 235B quantized model.
* If you go higher, you will OOM—even before context reaches the model's max.
---
## 4. Notes
* “Est. Total VRAM” is the minimum you’ll need for the model + context, but does not include OS, other processes, or toolbox/container overhead—leave a margin.
* For detailed methodology or custom scenarios, check the script source.
* Benchmark speed for large context sizes is often the real bottleneck—see `docs/benchmarks.md` for real throughput figures.
---
## 5. Related
* Main README section [Memory Planning & VRAM Estimator](../Readme#4--memory-planning--vram-estimator)
* [docs/benchmarks.md](benchmarks.md) for full speed/compat charts