SpacemiT K3 has 8× VLEN=1024 “AI” cores — we taught OpenBLAS about them and ran HPL

We spent some time on the Banana Pi BPI-SM10 (SpacemiT K3): 8× X100 (RVA23, RVV, VLEN=256) plus 8× A100 AI cores (VLEN=1024, IME2 matrix units). Both clusters speak RVV 1.0. The obvious question: can you just throw BLAS/HPL at the wide cores and go faster?

Short answer: no — not with stock VLEN=256 OpenBLAS kernels. Wide VLEN alone is not a free speedup.

Weird Linux bits first

A100s are online in sysfs but fenced from normal affinity. You need:

echo $$ > /proc/set_ai_thread

taskset -c 8-15 …

Also: don’t start on X100, let the dynamic linker / OpenBLAS cache VLEN=256 decisions, then migrate to A100. That path can SIGSEGV. Register the AI thread before exec.

HPL numbers (N=12000, NB=192, residuals PASSED)

Setup OpenBLAS GFLOP/s
X100 ×8 ZVL256B ~52–53
A100 ×8 ZVL256B (VLEN-safe) 15
A100 ×8 new ZVL1024B 35.9
8 X100 + 8 A100 ZVL256B + ZVL1024B 57.5

So A100 with stock-ish ZVL256B is ~3.5× slower than X100. Matching the tile to VLEN=1024 (RISCV64_ZVL1024B, DGEMM 16×8 from OpenBLAS’s generate_kernel.py) gets ~2.4× back on A100 — still behind X100, but now useful.

The fun part: once A100 isn’t a straggler, equal 8+8 mixed ranks beat X100-only (57.5 vs ~53). With the slow BLAS on A100, equal ranks were a trap (~27 GF); the old optimum was “mostly X100, few A100 ranks.”

Full L2/L3 OpenBLAS BLATs (S/D/C/Z) on A100 with the new target: all PASSED.

Takeaways

  1. K3 is heterogeneous. X100 remains the primary RVV compute path. A100’s width is really there for IME2, not “faster X60.”
  2. Kernel geometry must match VLEN. Porting ZVL256B kernels to VLEN=1024 without changing the tile wastes the machine (and some stock wide-load/vget patterns assume VLMAX@256).
  3. Heterogeneous MPI needs per-rank BLAS (we used FlexiBLAS) plus the set_ai_thread dance.
  4. Static TARGET=RISCV64_ZVL1024B must not run on X100 — use DYNAMIC_ARCH or separate backends.

Draft PR (kept draft on purpose — not proposing this as K3’s main path):
#6066

Board write-up + more tables:
https://www.opensolvers.com/boards/SM10.html

Patches: GitHub - opensolvers/benchmarks: Open scientific software & AI inference on real RISC-V hardware · GitHub (OpenBLAS/)

Board provided by Banana Pi. IME2 / GPU numbers still to come — this was the plain-RVV story only.

Happy to answer questions about the affinity trap, the vget bug, or the hetero HPL wrapper.