Benchmark · KubeMQ v3.6.16 vs Apache Kafka 4.3.1 · identical GCP hardware · 2026-10-06 / 2026-10-04

Same hardware. Same client. Which broker carries more, faster?

Three n2-standard-8 servers on local NVMe per broker, one OpenMessaging Benchmark client with byte-identical settings. Throughput ladder, latency at every rung, 100 KB messages, a leader kill with a loss verifier, and what each broker spends in CPU, memory and disk.

Run and published by KubeMQ, the vendor of one of the two systems. Everything needed to rerun it is in the repository: cluster creation, both deployments, the frozen workloads, the runner, the raw results. The KubeMQ half needs a licence key from KubeMQ. Known weaknesses: one run per cell; Kafka configuration tuned and reviewed only by us, the vendor, not by an outside Kafka engineer; the two brokers' acknowledgement semantics differ (see Rig); the two sides ran on different days and cluster incarnations (KubeMQ 2026-10-06, Kafka 2026-10-04), same hardware shape, baselines within 10 %.

Scorecard

Every number is one run (5 measured minutes; 15 for the 100k and kill cells). No repeats, so no spread is shown. There is no pass/fail verdict on this page: the numbers are reported as measured. "Reached" means the brokers delivered at least 99 % of the target rate with the producers on schedule; that is how the rate ladder decided when to stop climbing.
One run per cell. Bold = better by more than 10 %. No bold = within 10 % (a tie) or not comparable.
MeasureKubeMQ v3.6.16Apache Kafka 4.3.1Note
Throughput, 1 KB messages, 16 partitions, 3 replicas
Highest target rate reached, 128 KB client batches700k msg/s300k msg/sachieved ≥ 99 % of target and producers on schedule; 100k steps
Highest target rate reached, 512 KB client batches700k msg/s400k msg/ssame definition
Next rung after the top, 128 KB client800k target → 768k achieved; 96.0 % of target; producers 18.4 s behind schedule400k target → 400k achieved; publish-delay p99 838 mswhy each ladder ended
End-to-end latency p50 / p99, 1 KB, 128 KB client batches
at 100k msg/s (15 min)2 / 5 ms2 / 114 msp99 compared; a side that did not reach the target is shown but not ranked
at 300k msg/s4 / 13 ms3 / 274 msp99 compared; a side that did not reach the target is shown but not ranked
at 400k msg/s4 / 72 ms20 / 625 ms (publish-delay p99 838 ms)p99 compared; a side that did not reach the target is shown but not ranked
at 500k msg/s5 / 36 ms—p99 compared; a side that did not reach the target is shown but not ranked
at 600k msg/s6 / 32 ms—p99 compared; a side that did not reach the target is shown but not ranked
Publish latency p99 (producer ack), 1 KB, 128 KB client batches
at 100k msg/s3 ms114 ms
at 300k msg/s5 ms275 ms
at 400k msg/s6 ms631 ms
Large messages, 100 KB, 512 KB client batches
Highest target rate reached6,000 msg/s = 614 MB/s4,000 msg/s = 410 MB/spayload MB/s = rate × 0.1024
Achieved at 6,000 msg/s target (614 MB/s)6,004 msg/s = 615 MB/s4,694 msg/s = 481 MB/sdisk sequential-write ceiling on this rig ≈ 820 MB/s per broker; "extra" = run past that side's ladder end
Achieved at 8,000 msg/s target (819 MB/s)7,717 msg/s = 790 MB/s4,738 msg/s = 485 MB/s (extra)disk sequential-write ceiling on this rig ≈ 820 MB/s per broker; "extra" = run past that side's ladder end
Resilience: partition leader killed (SIGKILL) at 50k msg/s
Acknowledged messages lost00acked-ID ledger verifier beside the load
Longest interval with no acknowledgements0 s8 s
Killed server Ready again14 s23 s
All replicas back in sync94 s59 sKubeMQ: "All replicas in sync 94 s" is an upper bound: the in-sync check runs only after the restart proof, which finished 81 s after the kill
Resources at 300k msg/s, 1 KB, 128 KB client (highest rung both reached)
Broker CPU, cores of 7 (mean per broker)4.52.8
Broker memory working set (cgroup, includes page cache)2.5 GiB24.6 GiBKafka: 6 GiB heap + page cache; KubeMQ: Go heap + page cache
Disk written per broker332 MB/s314 MB/s
Bytes written per payload byte1.081.02write amplification incl. replication log
Network per broker (in + out)649 MB/s635 MB/s

Rate ladder, 1 KB messages

Fixed producer rate, 100k msg/s steps, 5 measured minutes per rung (15 at 100k). Each side climbed until a rung achieved less than 99 % of its target or the producers fell behind schedule (publish-delay p99 ≥ 100 ms); that rung is shown with a cross. Rungs marked "extra" were run beyond that point by owner decision or by a harness ordering defect and are kept as data.

KubeMQ v3.6.16Kafka 4.3.1dot = target reached, × = ladder ended here; label = e2e p99 ms
1101001,000 end-to-end p99, ms (log scale) 100k200k300k400k500k600k700k800k target rate, msg/s 114 232 274 625 5 7 13 72 36 32 34 640
128 KB client batches (cells A1 and B).
Ladder, 128 KB client batches (cells A1 and B). One run per rung.
TargetBrokerAchievedLast 60 sBacklog ΔPub-delay p99Publish p50 / p99 / p99.9E2E p50 / p99 / p99.9Disk MB/sCPU coresLimit seen
100kKubeMQ100.0k (100.0 %)99.9 %-2940.07 ms1.6 / 3 / 12 ms2 / 5 / 26 ms1143.8none
Kafka100.0k (100.0 %)99.9 %-5300.07 ms1.6 / 114 / 246 ms2 / 114 / 245 ms1042.0client window
200kKubeMQ extra200.2k (100.1 %)99.9 %-1160.07 ms1.8 / 4 / 12 ms3 / 7 / 209 ms2234.2none
Kafka extra200.2k (100.1 %)100.0 %620.07 ms1.9 / 233 / 353 ms2 / 232 / 351 ms2062.4client window
300kKubeMQ300.3k (100.1 %)99.9 %-2,6180.07 ms2.1 / 5 / 11 ms4 / 13 / 214 ms3324.5none
Kafka300.4k (100.1 %)99.9 %120.07 ms2.2 / 275 / 381 ms3 / 274 / 378 ms3142.8client window
400kKubeMQ400.3k (100.1 %)99.9 %1,0880.07 ms2.3 / 6 / 11 ms4 / 72 / 249 ms4404.8none
Kafka400.2k (100.0 %)99.6 %-63838 ms19.9 / 631 / 945 ms20 / 625 / 936 ms4193.1client window; publish-delay p99 838 ms
500kKubeMQ500.8k (100.2 %)100.0 %1,2080.07 ms2.6 / 9 / 16 ms5 / 36 / 251 ms5465.1none
Kafkanot run: this side's ladder had ended
600kKubeMQ600.5k (100.1 %)100.0 %-5020.06 ms3.2 / 15 / 25 ms6 / 32 / 264 ms6515.4none
Kafkanot run: this side's ladder had ended
700kKubeMQ700.4k (100.1 %)99.9 %150.06 ms4.0 / 22 / 33 ms9 / 34 / 54 ms7555.4disk
Kafkanot run: this side's ladder had ended
800kKubeMQ767.8k (96.0 %)95.9 %10,51218,378 ms152.8 / 472 / 861 ms238 / 640 / 960 ms8014.0disk; 96.0 % of target; producers 18.4 s behind schedule
Kafkanot run: this side's ladder had ended
1101001,00010,000 end-to-end p99, ms (log scale) 300k400k500k700k800k target rate, msg/s 387 1278 2154 89 1990
512 KB client batches (cell A2).
Ladder, 512 KB client batches (cell A2). One run per rung.
TargetBrokerAchievedLast 60 sBacklog ΔPub-delay p99Publish p50 / p99 / p99.9E2E p50 / p99 / p99.9Disk MB/sCPU coresLimit seen
300kKubeMQnot run: this side's ladder had ended
Kafka300.1k (100.0 %)100.1 %7960.07 ms2.7 / 383 / 717 ms3 / 387 / 717 ms3142.7client window
400kKubeMQnot run: this side's ladder had ended
Kafka400.5k (100.1 %)100.0 %-5840.07 ms88.3 / 1279 / 1599 ms92 / 1278 / 1599 ms4162.7client window
500kKubeMQnot run: this side's ladder had ended
Kafka481.0k (96.2 %)93.6 %46211,610 ms965.7 / 2155 / 2520 ms966 / 2154 / 2521 ms4982.8client window; 96.2 % of target; producers 11.6 s behind schedule
700kKubeMQ700.6k (100.1 %)99.9 %7,7410.06 ms4.7 / 35 / 98 ms14 / 89 / 223 ms7534.8disk
Kafkanot run: this side's ladder had ended
800kKubeMQ756.8k (94.6 %)94.4 %16,45123,738 ms658.5 / 1422 / 2171 ms813 / 1990 / 2640 ms7893.8disk; 94.6 % of target; producers 23.7 s behind schedule
Kafkanot run: this side's ladder had ended

Why is Kafka's tail long at light load?

Not explained by this run. At 100k msg/s Kafka's brokers were nearly idle (request-handler idle at least 98 %, GC pauses 18 to 46 ms per 10 s across three brokers, no old-generation collections) and Kafka's own Produce request timers showed a 99th percentile of 1 to 3 ms, yet the client measured a publish p99 above 50 ms in 51 of 90 ten-second windows, up to 252 ms. GC, log flushes and segment rolls do not line up with the slow windows. The delay therefore sits where this run has no instrument: in the client's accumulator and sender, or on the network before the broker reads the request. Kafka's exported percentiles are per request and have no maximum, so a few slow requests carrying many records could also hide under the broker's 99th. Analysis and per-window table: results/2026-10-04/notes/kafka-light-load.md.

100 KB messages

Bandwidth-bound workload with the 512 KB client (a 100 KB record cannot share a 128 KB batch). Disk sequential write on this rig is 820 MB/s per broker (fio). Both sides ran every rung by owner decision; rungs past the point where a side stopped reaching its target are marked "extra".

Cell C, one run per rung.
TargetPayload MB/sBrokerAchievedLast 60 sBacklog ΔPub-delay p99Publish p50 / p99 / p99.9E2E p50 / p99 / p99.9Disk MB/sCPU coresLimit seen
4,000410KubeMQ4.0k (100.1 %)99.9 %-630.10 ms2.5 / 7 / 12 ms5 / 13 / 21 ms4314.0none
Kafka4.0k (100.2 %)100.1 %120.10 ms108.3 / 1231 / 1548 ms107 / 1228 / 1545 ms4112.4client window
6,000614KubeMQ6.0k (100.1 %)99.9 %-230.11 ms3.6 / 21 / 34 ms8 / 37 / 67 ms6414.6none
Kafka4.7k (78.2 %)79.5 %-260,000 ms1061.5 / 2266 / 2732 ms1059 / 2262 / 2727 ms4792.5client window; 78.2 % of target; producers 60.0 s behind schedule
8,000819KubeMQ7.7k (96.5 %)95.9 %36615,419 ms613.4 / 1683 / 2156 ms764 / 1866 / 2286 ms7973.8disk; 96.5 % of target; producers 15.4 s behind schedule
Kafka extra4.7k (59.2 %)59.0 %060,000 ms1057.2 / 2277 / 2688 ms1055 / 2272 / 2684 ms4892.7client window; 59.2 % of target; producers 60.0 s behind schedule
10kKubeMQnot run: this side's ladder had ended
Kafka extra4.7k (46.7 %)46.2 %-560,000 ms1062.0 / 2391 / 2947 ms1060 / 2387 / 2948 ms4832.5client window; 46.7 % of target; producers 60.0 s behind schedule

Leader kill at 50k msg/s

At minute 5 of a 15-minute run the server leading the most partitions receives SIGKILL from the node and restarts on the same volume. A separate verifier publishes with an acked-ID ledger; every acked message must be read back. One kill per side.

Cell D.
KubeMQKafka
Acked messages lost (verifier verdict)0 of 3,900,000 (PASS)0 of 3,786,100 (PASS)
Longest interval with zero acks0 s8 s
Killed server Ready14 s23 s
All replicas in sync94 s59 s
Victim was also controller / control-shard leaderno — derived: shard-1 term stayed 2 across the kill; bench-0 follower before and after (kill/logs bench-0 04:45:29, 11:45:53)no — derived: restarted node rejoined as KRaft follower, epoch 1 unchanged, leader node 2 (kill/boot/bench-combined-0.log.txt 08:58:02)
Locator tie-break usednono
Kill proof (restart count +1, same volume inode, same image digest, survivors untouched)PASSPASS
Timeouts left at defaultRaft election (KubeMQ defaults)broker.session.timeout.ms 9 s

KubeMQ note: "All replicas in sync 94 s" is an upper bound: the in-sync check runs only after the restart proof, which finished 81 s after the kill. The kill landed at measured minute 6, not 5 (locator timing).

Resources

From 1-second Prometheus samples on the broker nodes during the measured window, per rung both sides ran, 128 KB client. The highlighted row is the highest rung both sides reached. "Cores per 100k" = sum of the three brokers' CPU ÷ carried rate. Memory is the container working set as the kernel accounts it, which includes page cache. Bold follows the tie rule.

Cell E, per broker unless stated.
RungBrokerCPU cores (of 7)Working setPage cache (node)Disk write MB/sDisk busyBytes written / payload byteNIC MB/sCores per 100k (cluster)$ per billion msgs
100kKubeMQ3.83.0 GiB27 GiB11415 %1.1123711.5$3.2
Kafka2.015.5 GiB24 GiB10412 %1.022165.9$3.2
200kKubeMQ extra4.22.2 GiB27 GiB22327 %1.094426.2$1.6
Kafka extra2.417.4 GiB24 GiB20625 %1.004223.6$1.6
300kKubeMQ4.52.5 GiB27 GiB33240 %1.086494.5$1.1
Kafka2.824.6 GiB24 GiB31438 %1.026352.8$1.1
400kKubeMQ4.82.8 GiB27 GiB44048 %1.078563.6$0.8
Kafka (target not reached)3.124.8 GiB24 GiB41951 %1.028412.3—
500kKubeMQ5.13.3 GiB27 GiB54655 %1.061,0613.1$0.6
600kKubeMQ5.43.8 GiB27 GiB65163 %1.061,2672.7$0.5
700kKubeMQ5.44.3 GiB26 GiB75567 %1.051,4652.3$0.5
800kKubeMQ (target not reached)4.06.5 GiB26 GiB80191 %1.021,6001.6—

$ per billion messages = hourly on-demand list price of that cluster's 3 broker nodes ($1.17/h) ÷ messages carried per hour, × 10⁹.

Rig and versions

Both clusters identical except the broker. Baselines per cluster incarnation, from results/2026-10-06/infra/.
Brokers3 × n2-standard-8 (Intel(R) Xeon(R) CPU @ 2.60GHz), 7 CPU / 26 GiB per broker, one per node
Disk2 × 375 GB local NVMe as RAID-0; fio sequential write 820 / 820 MB/s (Kafka / KubeMQ cluster), synced writes 12,071 / 11,922 per s
Networkiperf3 between brokers: 15.5 / 15.4 Gbit/s
Driver2 × n2-standard-8, OpenMessaging Benchmark (commit in omb/VERSION), stock Kafka driver, identical YAML both sides except bootstrap; 128 KB and 512 KB batch profiles
Kafka4.3.1, KRaft combined broker+controller (non-recommended for production, chosen for symmetry with KubeMQ's in-process Raft), Strimzi 1.2.0, image quay.io/strimzi/kafka@sha256:e90a1a74af4226f3ca4d1ebef3ab13bdb09754ae17ca4c1444f7fcbb0ca8ea9a, heap 6 GiB, tuned config in infra/kafka/kafka.yaml, RF 3, min ISR 2
KubeMQv3.6.16, image europe-docker.pkg.dev/kubemq/images/kubemq-next:v3.6.16@sha256:c9d54f1db925a3e792b5a11c2e0af9778fa1778c102eea70226fd25b1a4bed26, chart 3.5.0, shipped defaults (8 shards, leaders balanced, append-only log, default ack policy), GOMEMLIMIT 20 GiB
What an ack meansKafka: acks=all, the leader waits for every in-sync replica (normally 3 of 3) to have the write in the OS; no per-write fsync. KubeMQ: ack when 2 of 3 Raft members have the write in the OS; fsync in the background. These are not identical.

Method

Fixed-rate runs only (OMB's unthrottled mode understates one side). The ladder climbed in 100k steps until a side achieved less than 99 % of target or its producers fell behind schedule; later rungs run by owner decision are marked "extra". No pass/fail verdict is applied; latency and resource figures are reported as measured at every rung. One run per cell; a run with an infrastructure fault is VOID and rerun once (one Kafka kill attempt was voided for a harness defect and rerun). One discarded warm-up after every deploy; topic deleted and disk hygiene checked between runs. Kafka: tuned in an earlier session (tuning/kafka.md); KubeMQ: shipped defaults. Not tested: fan-out, soak, backlog drain, many partitions, strict fsync, repeats. Full plan: docs/PLAN.md; review: docs/REVIEW.md.

Reproduce

echo "$KEY" > .license.txt
terraform -chdir=infra/terraform init
scripts/build-images.sh
scripts/up.sh kafka ; sleep 120 ; scripts/up.sh kubemq
caffeinate -i scripts/matrix.sh --slot a    # terminal 1
caffeinate -i scripts/matrix.sh --slot b    # terminal 2
scripts/report.py results/2026-10-06
scripts/down.sh --slot a ; scripts/down.sh --slot b

Raw OMB JSON, 1 s Prometheus exports, verifier ledgers, fio and iperf baselines, image digests: results/2026-10-06/.

Generated by scripts/report.py from results/2026-10-06 on 2026-10-06 14:56 UTC.

Run it on your own numbers

Every script, workload and raw result is in the repository. Bring up the same two clusters, run the matrix, and compare against your own traffic shape.

Start a KubeMQ trial