Scorecard
| Measure | KubeMQ v3.6.16 | Apache Kafka 4.3.1 | Note |
|---|---|---|---|
| Throughput, 1 KB messages, 16 partitions, 3 replicas | |||
| Highest target rate reached, 128 KB client batches | 700k msg/s | 300k msg/s | achieved ≥ 99 % of target and producers on schedule; 100k steps |
| Highest target rate reached, 512 KB client batches | 700k msg/s | 400k msg/s | same definition |
| Next rung after the top, 128 KB client | 800k target → 768k achieved; 96.0 % of target; producers 18.4 s behind schedule | 400k target → 400k achieved; publish-delay p99 838 ms | why each ladder ended |
| End-to-end latency p50 / p99, 1 KB, 128 KB client batches | |||
| at 100k msg/s (15 min) | 2 / 5 ms | 2 / 114 ms | p99 compared; a side that did not reach the target is shown but not ranked |
| at 300k msg/s | 4 / 13 ms | 3 / 274 ms | p99 compared; a side that did not reach the target is shown but not ranked |
| at 400k msg/s | 4 / 72 ms | 20 / 625 ms (publish-delay p99 838 ms) | p99 compared; a side that did not reach the target is shown but not ranked |
| at 500k msg/s | 5 / 36 ms | — | p99 compared; a side that did not reach the target is shown but not ranked |
| at 600k msg/s | 6 / 32 ms | — | p99 compared; a side that did not reach the target is shown but not ranked |
| Publish latency p99 (producer ack), 1 KB, 128 KB client batches | |||
| at 100k msg/s | 3 ms | 114 ms | |
| at 300k msg/s | 5 ms | 275 ms | |
| at 400k msg/s | 6 ms | 631 ms | |
| Large messages, 100 KB, 512 KB client batches | |||
| Highest target rate reached | 6,000 msg/s = 614 MB/s | 4,000 msg/s = 410 MB/s | payload MB/s = rate × 0.1024 |
| Achieved at 6,000 msg/s target (614 MB/s) | 6,004 msg/s = 615 MB/s | 4,694 msg/s = 481 MB/s | disk sequential-write ceiling on this rig ≈ 820 MB/s per broker; "extra" = run past that side's ladder end |
| Achieved at 8,000 msg/s target (819 MB/s) | 7,717 msg/s = 790 MB/s | 4,738 msg/s = 485 MB/s (extra) | disk sequential-write ceiling on this rig ≈ 820 MB/s per broker; "extra" = run past that side's ladder end |
| Resilience: partition leader killed (SIGKILL) at 50k msg/s | |||
| Acknowledged messages lost | 0 | 0 | acked-ID ledger verifier beside the load |
| Longest interval with no acknowledgements | 0 s | 8 s | |
| Killed server Ready again | 14 s | 23 s | |
| All replicas back in sync | 94 s | 59 s | KubeMQ: "All replicas in sync 94 s" is an upper bound: the in-sync check runs only after the restart proof, which finished 81 s after the kill |
| Resources at 300k msg/s, 1 KB, 128 KB client (highest rung both reached) | |||
| Broker CPU, cores of 7 (mean per broker) | 4.5 | 2.8 | |
| Broker memory working set (cgroup, includes page cache) | 2.5 GiB | 24.6 GiB | Kafka: 6 GiB heap + page cache; KubeMQ: Go heap + page cache |
| Disk written per broker | 332 MB/s | 314 MB/s | |
| Bytes written per payload byte | 1.08 | 1.02 | write amplification incl. replication log |
| Network per broker (in + out) | 649 MB/s | 635 MB/s | |
Rate ladder, 1 KB messages
Fixed producer rate, 100k msg/s steps, 5 measured minutes per rung (15 at 100k). Each side climbed until a rung achieved less than 99 % of its target or the producers fell behind schedule (publish-delay p99 ≥ 100 ms); that rung is shown with a cross. Rungs marked "extra" were run beyond that point by owner decision or by a harness ordering defect and are kept as data.
| Target | Broker | Achieved | Last 60 s | Backlog Δ | Pub-delay p99 | Publish p50 / p99 / p99.9 | E2E p50 / p99 / p99.9 | Disk MB/s | CPU cores | Limit seen |
|---|---|---|---|---|---|---|---|---|---|---|
| 100k | KubeMQ | 100.0k (100.0 %) | 99.9 % | -294 | 0.07 ms | 1.6 / 3 / 12 ms | 2 / 5 / 26 ms | 114 | 3.8 | none |
| Kafka | 100.0k (100.0 %) | 99.9 % | -530 | 0.07 ms | 1.6 / 114 / 246 ms | 2 / 114 / 245 ms | 104 | 2.0 | client window | |
| 200k | KubeMQ extra | 200.2k (100.1 %) | 99.9 % | -116 | 0.07 ms | 1.8 / 4 / 12 ms | 3 / 7 / 209 ms | 223 | 4.2 | none |
| Kafka extra | 200.2k (100.1 %) | 100.0 % | 62 | 0.07 ms | 1.9 / 233 / 353 ms | 2 / 232 / 351 ms | 206 | 2.4 | client window | |
| 300k | KubeMQ | 300.3k (100.1 %) | 99.9 % | -2,618 | 0.07 ms | 2.1 / 5 / 11 ms | 4 / 13 / 214 ms | 332 | 4.5 | none |
| Kafka | 300.4k (100.1 %) | 99.9 % | 12 | 0.07 ms | 2.2 / 275 / 381 ms | 3 / 274 / 378 ms | 314 | 2.8 | client window | |
| 400k | KubeMQ | 400.3k (100.1 %) | 99.9 % | 1,088 | 0.07 ms | 2.3 / 6 / 11 ms | 4 / 72 / 249 ms | 440 | 4.8 | none |
| Kafka | 400.2k (100.0 %) | 99.6 % | -63 | 838 ms | 19.9 / 631 / 945 ms | 20 / 625 / 936 ms | 419 | 3.1 | client window; publish-delay p99 838 ms | |
| 500k | KubeMQ | 500.8k (100.2 %) | 100.0 % | 1,208 | 0.07 ms | 2.6 / 9 / 16 ms | 5 / 36 / 251 ms | 546 | 5.1 | none |
| Kafka | not run: this side's ladder had ended | |||||||||
| 600k | KubeMQ | 600.5k (100.1 %) | 100.0 % | -502 | 0.06 ms | 3.2 / 15 / 25 ms | 6 / 32 / 264 ms | 651 | 5.4 | none |
| Kafka | not run: this side's ladder had ended | |||||||||
| 700k | KubeMQ | 700.4k (100.1 %) | 99.9 % | 15 | 0.06 ms | 4.0 / 22 / 33 ms | 9 / 34 / 54 ms | 755 | 5.4 | disk |
| Kafka | not run: this side's ladder had ended | |||||||||
| 800k | KubeMQ | 767.8k (96.0 %) | 95.9 % | 10,512 | 18,378 ms | 152.8 / 472 / 861 ms | 238 / 640 / 960 ms | 801 | 4.0 | disk; 96.0 % of target; producers 18.4 s behind schedule |
| Kafka | not run: this side's ladder had ended | |||||||||
| Target | Broker | Achieved | Last 60 s | Backlog Δ | Pub-delay p99 | Publish p50 / p99 / p99.9 | E2E p50 / p99 / p99.9 | Disk MB/s | CPU cores | Limit seen |
|---|---|---|---|---|---|---|---|---|---|---|
| 300k | KubeMQ | not run: this side's ladder had ended | ||||||||
| Kafka | 300.1k (100.0 %) | 100.1 % | 796 | 0.07 ms | 2.7 / 383 / 717 ms | 3 / 387 / 717 ms | 314 | 2.7 | client window | |
| 400k | KubeMQ | not run: this side's ladder had ended | ||||||||
| Kafka | 400.5k (100.1 %) | 100.0 % | -584 | 0.07 ms | 88.3 / 1279 / 1599 ms | 92 / 1278 / 1599 ms | 416 | 2.7 | client window | |
| 500k | KubeMQ | not run: this side's ladder had ended | ||||||||
| Kafka | 481.0k (96.2 %) | 93.6 % | 462 | 11,610 ms | 965.7 / 2155 / 2520 ms | 966 / 2154 / 2521 ms | 498 | 2.8 | client window; 96.2 % of target; producers 11.6 s behind schedule | |
| 700k | KubeMQ | 700.6k (100.1 %) | 99.9 % | 7,741 | 0.06 ms | 4.7 / 35 / 98 ms | 14 / 89 / 223 ms | 753 | 4.8 | disk |
| Kafka | not run: this side's ladder had ended | |||||||||
| 800k | KubeMQ | 756.8k (94.6 %) | 94.4 % | 16,451 | 23,738 ms | 658.5 / 1422 / 2171 ms | 813 / 1990 / 2640 ms | 789 | 3.8 | disk; 94.6 % of target; producers 23.7 s behind schedule |
| Kafka | not run: this side's ladder had ended | |||||||||
Why is Kafka's tail long at light load?
results/2026-10-04/notes/kafka-light-load.md.100 KB messages
Bandwidth-bound workload with the 512 KB client (a 100 KB record cannot share a 128 KB batch). Disk sequential write on this rig is 820 MB/s per broker (fio). Both sides ran every rung by owner decision; rungs past the point where a side stopped reaching its target are marked "extra".
| Target | Payload MB/s | Broker | Achieved | Last 60 s | Backlog Δ | Pub-delay p99 | Publish p50 / p99 / p99.9 | E2E p50 / p99 / p99.9 | Disk MB/s | CPU cores | Limit seen | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4,000 | 410 | KubeMQ | 4.0k (100.1 %) | 99.9 % | -63 | 0.10 ms | 2.5 / 7 / 12 ms | 5 / 13 / 21 ms | 431 | 4.0 | none | |
| Kafka | 4.0k (100.2 %) | 100.1 % | 12 | 0.10 ms | 108.3 / 1231 / 1548 ms | 107 / 1228 / 1545 ms | 411 | 2.4 | client window | |||
| 6,000 | 614 | KubeMQ | 6.0k (100.1 %) | 99.9 % | -23 | 0.11 ms | 3.6 / 21 / 34 ms | 8 / 37 / 67 ms | 641 | 4.6 | none | |
| Kafka | 4.7k (78.2 %) | 79.5 % | -2 | 60,000 ms | 1061.5 / 2266 / 2732 ms | 1059 / 2262 / 2727 ms | 479 | 2.5 | client window; 78.2 % of target; producers 60.0 s behind schedule | |||
| 8,000 | 819 | KubeMQ | 7.7k (96.5 %) | 95.9 % | 366 | 15,419 ms | 613.4 / 1683 / 2156 ms | 764 / 1866 / 2286 ms | 797 | 3.8 | disk; 96.5 % of target; producers 15.4 s behind schedule | |
| Kafka extra | 4.7k (59.2 %) | 59.0 % | 0 | 60,000 ms | 1057.2 / 2277 / 2688 ms | 1055 / 2272 / 2684 ms | 489 | 2.7 | client window; 59.2 % of target; producers 60.0 s behind schedule | |||
| 10k | KubeMQ | not run: this side's ladder had ended | ||||||||||
| Kafka extra | 4.7k (46.7 %) | 46.2 % | -5 | 60,000 ms | 1062.0 / 2391 / 2947 ms | 1060 / 2387 / 2948 ms | 483 | 2.5 | client window; 46.7 % of target; producers 60.0 s behind schedule | |||
Leader kill at 50k msg/s
At minute 5 of a 15-minute run the server leading the most partitions receives SIGKILL from the node and restarts on the same volume. A separate verifier publishes with an acked-ID ledger; every acked message must be read back. One kill per side.
| KubeMQ | Kafka | |
|---|---|---|
| Acked messages lost (verifier verdict) | 0 of 3,900,000 (PASS) | 0 of 3,786,100 (PASS) |
| Longest interval with zero acks | 0 s | 8 s |
| Killed server Ready | 14 s | 23 s |
| All replicas in sync | 94 s | 59 s |
| Victim was also controller / control-shard leader | no — derived: shard-1 term stayed 2 across the kill; bench-0 follower before and after (kill/logs bench-0 04:45:29, 11:45:53) | no — derived: restarted node rejoined as KRaft follower, epoch 1 unchanged, leader node 2 (kill/boot/bench-combined-0.log.txt 08:58:02) |
| Locator tie-break used | no | no |
| Kill proof (restart count +1, same volume inode, same image digest, survivors untouched) | PASS | PASS |
| Timeouts left at default | Raft election (KubeMQ defaults) | broker.session.timeout.ms 9 s |
KubeMQ note: "All replicas in sync 94 s" is an upper bound: the in-sync check runs only after the restart proof, which finished 81 s after the kill. The kill landed at measured minute 6, not 5 (locator timing).
Resources
From 1-second Prometheus samples on the broker nodes during the measured window, per rung both sides ran, 128 KB client. The highlighted row is the highest rung both sides reached. "Cores per 100k" = sum of the three brokers' CPU ÷ carried rate. Memory is the container working set as the kernel accounts it, which includes page cache. Bold follows the tie rule.
| Rung | Broker | CPU cores (of 7) | Working set | Page cache (node) | Disk write MB/s | Disk busy | Bytes written / payload byte | NIC MB/s | Cores per 100k (cluster) | $ per billion msgs |
|---|---|---|---|---|---|---|---|---|---|---|
| 100k | KubeMQ | 3.8 | 3.0 GiB | 27 GiB | 114 | 15 % | 1.11 | 237 | 11.5 | $3.2 |
| Kafka | 2.0 | 15.5 GiB | 24 GiB | 104 | 12 % | 1.02 | 216 | 5.9 | $3.2 | |
| 200k | KubeMQ extra | 4.2 | 2.2 GiB | 27 GiB | 223 | 27 % | 1.09 | 442 | 6.2 | $1.6 |
| Kafka extra | 2.4 | 17.4 GiB | 24 GiB | 206 | 25 % | 1.00 | 422 | 3.6 | $1.6 | |
| 300k | KubeMQ | 4.5 | 2.5 GiB | 27 GiB | 332 | 40 % | 1.08 | 649 | 4.5 | $1.1 |
| Kafka | 2.8 | 24.6 GiB | 24 GiB | 314 | 38 % | 1.02 | 635 | 2.8 | $1.1 | |
| 400k | KubeMQ | 4.8 | 2.8 GiB | 27 GiB | 440 | 48 % | 1.07 | 856 | 3.6 | $0.8 |
| Kafka (target not reached) | 3.1 | 24.8 GiB | 24 GiB | 419 | 51 % | 1.02 | 841 | 2.3 | — | |
| 500k | KubeMQ | 5.1 | 3.3 GiB | 27 GiB | 546 | 55 % | 1.06 | 1,061 | 3.1 | $0.6 |
| 600k | KubeMQ | 5.4 | 3.8 GiB | 27 GiB | 651 | 63 % | 1.06 | 1,267 | 2.7 | $0.5 |
| 700k | KubeMQ | 5.4 | 4.3 GiB | 26 GiB | 755 | 67 % | 1.05 | 1,465 | 2.3 | $0.5 |
| 800k | KubeMQ (target not reached) | 4.0 | 6.5 GiB | 26 GiB | 801 | 91 % | 1.02 | 1,600 | 1.6 | — |
$ per billion messages = hourly on-demand list price of that cluster's 3 broker nodes ($1.17/h) ÷ messages carried per hour, × 10⁹.
Rig and versions
| Brokers | 3 × n2-standard-8 (Intel(R) Xeon(R) CPU @ 2.60GHz), 7 CPU / 26 GiB per broker, one per node |
| Disk | 2 × 375 GB local NVMe as RAID-0; fio sequential write 820 / 820 MB/s (Kafka / KubeMQ cluster), synced writes 12,071 / 11,922 per s |
| Network | iperf3 between brokers: 15.5 / 15.4 Gbit/s |
| Driver | 2 × n2-standard-8, OpenMessaging Benchmark (commit in omb/VERSION), stock Kafka driver, identical YAML both sides except bootstrap; 128 KB and 512 KB batch profiles |
| Kafka | 4.3.1, KRaft combined broker+controller (non-recommended for production, chosen for symmetry with KubeMQ's in-process Raft), Strimzi 1.2.0, image quay.io/strimzi/kafka@sha256:e90a1a74af4226f3ca4d1ebef3ab13bdb09754ae17ca4c1444f7fcbb0ca8ea9a, heap 6 GiB, tuned config in infra/kafka/kafka.yaml, RF 3, min ISR 2 |
| KubeMQ | v3.6.16, image europe-docker.pkg.dev/kubemq/images/kubemq-next:v3.6.16@sha256:c9d54f1db925a3e792b5a11c2e0af9778fa1778c102eea70226fd25b1a4bed26, chart 3.5.0, shipped defaults (8 shards, leaders balanced, append-only log, default ack policy), GOMEMLIMIT 20 GiB |
| What an ack means | Kafka: acks=all, the leader waits for every in-sync replica (normally 3 of 3) to have the write in the OS; no per-write fsync. KubeMQ: ack when 2 of 3 Raft members have the write in the OS; fsync in the background. These are not identical. |
Method
Fixed-rate runs only (OMB's unthrottled mode understates one side). The ladder climbed in 100k steps until a side achieved less than 99 % of target or its producers fell behind schedule; later rungs run by owner decision are marked "extra". No pass/fail verdict is applied; latency and resource figures are reported as measured at every rung. One run per cell; a run with an infrastructure fault is VOID and rerun once (one Kafka kill attempt was voided for a harness defect and rerun). One discarded warm-up after every deploy; topic deleted and disk hygiene checked between runs. Kafka: tuned in an earlier session (tuning/kafka.md); KubeMQ: shipped defaults. Not tested: fan-out, soak, backlog drain, many partitions, strict fsync, repeats. Full plan: docs/PLAN.md; review: docs/REVIEW.md.
Reproduce
echo "$KEY" > .license.txt terraform -chdir=infra/terraform init scripts/build-images.sh scripts/up.sh kafka ; sleep 120 ; scripts/up.sh kubemq caffeinate -i scripts/matrix.sh --slot a # terminal 1 caffeinate -i scripts/matrix.sh --slot b # terminal 2 scripts/report.py results/2026-10-06 scripts/down.sh --slot a ; scripts/down.sh --slot b
Raw OMB JSON, 1 s Prometheus exports, verifier ledgers, fio and iperf baselines, image digests: results/2026-10-06/.
Generated by scripts/report.py from results/2026-10-06 on 2026-10-06 14:56 UTC.
Run it on your own numbers
Every script, workload and raw result is in the repository. Bring up the same two clusters, run the matrix, and compare against your own traffic shape.