状态: 已发布数据所有权: 研究院所有许可: CC BY 4.0
AWS L4(24GB)全网格基准
在租用的 AWS L4 24GB 实例上,对各引擎的一次性能力快照。
用途
一次能力快照,而非经调优或重复的基准。
来源
在租用 L4 上的一次性基准会话。
构成
可下载 CSV——一次性的行:(engine, task, metric, value, peak_vram_gb, note):自主 text-CUDA 的 Qwen3-30B-A3B q4 全 48 层常驻解码结果(约 23 tok/s,19.73 GB),以及本次自主但无法运行的图像/视频构建。
已知偏差
- text-cuda 解码从主机内存流式加载权重(非常驻 GPU),确实慢于常驻路径。
局限
- 单次运行,无方差。
- 图像/视频构建虽自主,但因缺可移植权重布局而无法运行。
访问条件
所测快照可在下方以 CC-BY-4.0 下载。
| engine | task | metric | value | peak_vram_gb | note |
|---|---|---|---|---|---|
| text-cuda-moe | Qwen3-30B-A3B q4 full-48 resident decode | tok_per_s | 23 | 19.73 | sovereign (ldd = libcudart only; no cuBLAS/cuDNN/cuTLASS); weights 18.375 GB; parity logits_cosine=1.0 argmax 12/12 |
| image-cuda | SDXL 1024px full-step | status | not-run | sovereign build but unrunnable on this run pending portable weight layouts | |
| video-cuda | Wan2.1 full-frame | status | not-run | sovereign build but unrunnable on this run pending portable weight layouts |
