跳到主要内容
Octopus Research Institute
状态: 已发布数据所有权: 研究院所有许可: CC BY 4.0

AWS L4(24GB)全网格基准

在租用的 AWS L4 24GB 实例上,对各引擎的一次性能力快照。

用途

一次能力快照,而非经调优或重复的基准。

来源

在租用 L4 上的一次性基准会话。

构成

可下载 CSV——一次性的行:(engine, task, metric, value, peak_vram_gb, note):自主 text-CUDA 的 Qwen3-30B-A3B q4 全 48 层常驻解码结果(约 23 tok/s,19.73 GB),以及本次自主但无法运行的图像/视频构建。

已知偏差

  • text-cuda 解码从主机内存流式加载权重(非常驻 GPU),确实慢于常驻路径。

局限

  • 单次运行,无方差。
  • 图像/视频构建虽自主,但因缺可移植权重布局而无法运行。

访问条件

所测快照可在下方以 CC-BY-4.0 下载。

数据预览
enginetaskmetricvaluepeak_vram_gbnote
text-cuda-moeQwen3-30B-A3B q4 full-48 resident decodetok_per_s2319.73sovereign (ldd = libcudart only; no cuBLAS/cuDNN/cuTLASS); weights 18.375 GB; parity logits_cosine=1.0 argmax 12/12
image-cudaSDXL 1024px full-stepstatusnot-runsovereign build but unrunnable on this run pending portable weight layouts
video-cudaWan2.1 full-framestatusnot-runsovereign build but unrunnable on this run pending portable weight layouts
共 3 行,显示前 3 行。
下载 · CC BY 4.0