Performance
minc compared with clang and MSVC on the box3d benchmark suite.
box3d is a 3D rigid-body physics engine: fifty source files of mostly float math and SIMD C, with a benchmark suite of eleven scenes. The C compilers build the upstream source. The minc compiler builds box3d-minc, generated from the same source by a C-minc-transpiler written in minc. All builds emit similar 128-bit SIMD code.
Wall-clock milliseconds per scene, single-threaded, lower is better.
Windows x64
| scene | minc | clang | MSVC | vs clang | vs MSVC |
|---|---|---|---|---|---|
| many_pyramids | 1,841.4 | 1,940.0 | 2,056.2 | 0.94× | 0.89× |
| large_pyramid | 1,793.2 | 1,903.1 | 1,987.5 | 0.95× | 0.91× |
| joint_grid | 1,286.7 | 1,160.1 | 1,448.2 | 1.11× | 0.89× |
| washer | 24,566.3 | 21,920.3 | 23,352.8 | 1.13× | 1.06× |
| rain | 2,231.0 | 1,970.4 | 2,121.6 | 1.13× | 1.05× |
| large_world | 8.6 | 7.5 | 8.0 | 1.16× | 1.08× |
| junkyard | 18,043.0 | 15,122.4 | 16,439.4 | 1.19× | 1.10× |
| convex_pile | 17,283.0 | 13,181.7 | 13,809.4 | 1.31× | 1.25× |
| trees100 | 240.5 | 183.2 | 213.6 | 1.32× | 1.13× |
| trees25 | 940.5 | 676.5 | 840.9 | 1.39× | 1.12× |
| trees50 | 382.6 | 273.1 | 340.3 | 1.40× | 1.12× |
| geometric mean | — | — | — | 1.18× | 1.05× |
Across the suite minc is 1.18× clang and
1.05× MSVC. It matches or beats clang on the stacking
scenes, beats MSVC on three of the eleven, and trails most on the
mesh-collision (trees) scenes.
With the solver spread across worker threads the three compilers scale together. Scaling is against the same compiler's own single-worker total.
| workers | minc total | clang total | MSVC total | vs clang | vs MSVC | minc scaling | clang scaling | MSVC scaling |
|---|---|---|---|---|---|---|---|---|
| 1 | 68,617 | 58,338 | 62,618 | 1.176× | 1.049× | — | — | — |
| 2 | 36,308 | 31,267 | 33,108 | 1.168× | 1.054× | 1.89× | 1.87× | 1.89× |
| 4 | 20,489 | 17,742 | 18,682 | 1.167× | 1.060× | 3.35× | 3.29× | 3.35× |
| 8 | 14,560 | 13,246 | 13,627 | 1.132× | 1.046× | 4.71× | 4.40× | 4.60× |
macOS arm64
The same eleven scenes on an Apple M1, against Apple clang.
| scene | minc | clang | vs clang |
|---|---|---|---|
| many_pyramids | 2,176.2 | 2,077.3 | 1.05× |
| large_pyramid | 2,005.1 | 1,872.0 | 1.07× |
| washer | 26,143.7 | 24,268.5 | 1.08× |
| junkyard | 16,886.7 | 14,904.5 | 1.13× |
| large_world | 7.4 | 6.3 | 1.18× |
| rain | 2,510.4 | 2,063.6 | 1.22× |
| trees100 | 244.1 | 198.4 | 1.23× |
| joint_grid | 1,255.2 | 991.8 | 1.26× |
| convex_pile | 18,916.6 | 14,741.0 | 1.28× |
| trees50 | 389.2 | 290.5 | 1.34× |
| trees25 | 978.5 | 699.7 | 1.40× |
| geometric mean | — | — | 1.20× |
minc is 1.20× Apple clang, matching the Windows figure, and the scene ordering is also similar.
At 2, 4 and 8 workers the ratio stays within 0.02× of the single-worker figure, and minc's thread pool scales like clang's. Totals sum all eleven scenes; scaling is against the same compiler's own single-worker total.
| workers | minc total | clang total | ratio | geomean | minc scaling | clang scaling |
|---|---|---|---|---|---|---|
| 1 | 71,513 | 62,114 | 1.15× | 1.198× | — | — |
| 2 | 38,291 | 33,457 | 1.14× | 1.188× | 1.87× | 1.86× |
| 4 | 21,879 | 19,130 | 1.14× | 1.162× | 3.27× | 3.25× |
| 8 | 20,312 | 18,010 | 1.13× | 1.140× | 3.52× | 3.45× |
WebAssembly
The same eleven scenes compiled to WebAssembly and run in Node on the
Windows machine: minc builds the published modules with
--target wasm, emcc builds the upstream C from the same box3d
pin. Both emit 128-bit wasm SIMD. Single worker.
| scene | minc | emcc | vs emcc |
|---|---|---|---|
| many_pyramids | 2,907.1 | 2,683.8 | 1.08× |
| large_pyramid | 2,805.3 | 2,521.7 | 1.11× |
| joint_grid | 1,997.3 | 1,674.9 | 1.20× |
| large_world | 15.2 | 12.2 | 1.25× |
| rain | 3,637.4 | 2,915.4 | 1.26× |
| junkyard | 29,297.7 | 23,100.3 | 1.27× |
| washer | 40,476.7 | 30,718.2 | 1.32× |
| convex_pile | 26,881.7 | 19,954.2 | 1.35× |
| trees25 | 1,345.3 | 899.7 | 1.50× |
| trees100 | 356.7 | 238.2 | 1.50× |
| trees50 | 571.7 | 367.8 | 1.55× |
| geometric mean | — | — | 1.30× |
minc is 1.30× emcc on wasm, with the same shape as the native tables: stacking scenes closest, mesh collision the tail. emcc goes through LLVM, so this measures minc's wasm backend against the same optimizer clang uses natively.
Test setup
| Windows machine | AMD Ryzen 9 5900X, 12 cores / 24 threads, 32 GB, Windows 11, stock power plan, idle |
| macOS machine | Apple M1, 16 GB, macOS 14.8.1, AC power, idle |
| minc | 0.9.12, default flags (bounds checking on) |
| MSVC | 19.44.35227 (Visual Studio 2022, 17.14.32), /O2 /arch:AVX2 /fp:contract /std:c17 /DNDEBUG |
| clang-cl | 19.1.5, target x86_64-pc-windows-msvc (bundled with Visual Studio 2022), /O2 /arch:AVX2 /std:c17 /DNDEBUG |
| Apple clang | 16.0.0 (clang-1600.0.26.6, Xcode), -O2 -DNDEBUG |
| emcc | 4.0.19, -O2 -msimd128 -msse2 -std=gnu17 -DNDEBUG, run in Node v22.16.0 |
| box3d | upstream 2386141. minc builds the published box3d-minc module; the C compilers build the upstream sources directly |
| Date | 2026-08-17 (Windows, wasm), 2026-08-16 (macOS) |
Notes
- Method. Every binary runs a given scene back-to-back with the others, so a comparison pair is seconds apart. Each binary and scene gets one discarded warm-up run, then three timed rounds. The millisecond columns show the fastest round. The ratio columns show the median of the three per-round paired ratios, which cancels machine drift. Dividing the millisecond columns gives a slightly different number; the two methods agree to within 0.03× on every scene. On wasm each sample additionally takes the fastest of four in-process runs, the upstream driver's own default; a fresh process's first run carries V8 tier-up, which reads as a 2× slowdown on the shortest scene and is compile time, not generated code. Run-to-run spread at the 90th percentile is 2.1% on Windows, 1.0% on macOS and 3.4% on wasm. Ratios within a few percent of 1.00× are parity.
/fp:contractis required for a fair MSVC comparison. minc fuses multiply-add unconditionally and clang contracts by default. MSVC keeps multiply and add separate under/O2 /arch:AVX2alone, and measures about 10% slower without/fp:contract.- minc emits no automatic 256-bit SIMD. The x64 references use
ymmmostly as a copy width: about 1,900 ymm instructions in the clang binary and 2,200 in MSVC, dominated by 32-bytevmovups, against zero in minc. clang also vectorizes a few float loops to 256-bit. MSVC keeps float math at 128-bit. minc's vector types currently cover explicit 256-bit SIMD only. Box3d's own 128-bit SIMD path, used on both platforms, compiles in full. - Bounds checking is enabled in minc. The C columns run unchecked.
- Single-threaded. box3d can spread the solver across workers, but the per-scene tables use one worker. The scaling tables cover 2, 4 and 8 workers on both machines; the ratios hold or narrow as workers are added.
- This is one benchmark suite, and work in progress. Physics is float- and branch-heavy and leans on struct-by-value math. Other workloads will show different ratios.