We spent some time on the Banana Pi BPI-SM10 (SpacemiT K3): 8× X100 (RVA23, RVV, VLEN=256) plus 8× A100 AI cores (VLEN=1024, IME2 matrix units). Both clusters speak RVV 1.0. The obvious question: can you just throw BLAS/HPL at the wide cores and go faster?
Short answer: no — not with stock VLEN=256 OpenBLAS kernels. Wide VLEN alone is not a free speedup.
Weird Linux bits first
A100s are online in sysfs but fenced from normal affinity. You need:
echo $$ > /proc/set_ai_thread
taskset -c 8-15 …
Also: don’t start on X100, let the dynamic linker / OpenBLAS cache VLEN=256 decisions, then migrate to A100. That path can SIGSEGV. Register the AI thread before exec.
HPL numbers (N=12000, NB=192, residuals PASSED)
| Setup | OpenBLAS | GFLOP/s |
|---|---|---|
| X100 ×8 | ZVL256B | ~52–53 |
| A100 ×8 | ZVL256B (VLEN-safe) | 15 |
| A100 ×8 | new ZVL1024B | 35.9 |
| 8 X100 + 8 A100 | ZVL256B + ZVL1024B | 57.5 |
So A100 with stock-ish ZVL256B is ~3.5× slower than X100. Matching the tile to VLEN=1024 (RISCV64_ZVL1024B, DGEMM 16×8 from OpenBLAS’s generate_kernel.py) gets ~2.4× back on A100 — still behind X100, but now useful.
The fun part: once A100 isn’t a straggler, equal 8+8 mixed ranks beat X100-only (57.5 vs ~53). With the slow BLAS on A100, equal ranks were a trap (~27 GF); the old optimum was “mostly X100, few A100 ranks.”
Full L2/L3 OpenBLAS BLATs (S/D/C/Z) on A100 with the new target: all PASSED.
Takeaways
- K3 is heterogeneous. X100 remains the primary RVV compute path. A100’s width is really there for IME2, not “faster X60.”
- Kernel geometry must match VLEN. Porting ZVL256B kernels to VLEN=1024 without changing the tile wastes the machine (and some stock wide-load/
vgetpatterns assume VLMAX@256). - Heterogeneous MPI needs per-rank BLAS (we used FlexiBLAS) plus the
set_ai_threaddance. - Static
TARGET=RISCV64_ZVL1024Bmust not run on X100 — use DYNAMIC_ARCH or separate backends.
Draft PR (kept draft on purpose — not proposing this as K3’s main path):
#6066
Board write-up + more tables:
https://www.opensolvers.com/boards/SM10.html
Patches: GitHub - opensolvers/benchmarks: Open scientific software & AI inference on real RISC-V hardware · GitHub (OpenBLAS/)
Board provided by Banana Pi. IME2 / GPU numbers still to come — this was the plain-RVV story only.
Happy to answer questions about the affinity trap, the vget bug, or the hetero HPL wrapper.