Skip to content

quicX Performance Baseline

Platform: macOS ARM64 (Apple Silicon M3 Pro, 14 cores)

Build: CMake RelWithDebInfo, Clang, -O2 -g

Framework: Google Benchmark v1.8.3


BenchmarkTimeThroughputNote
TlsCtxCreation~16 ns—TLS context object creation (lightweight)

1.2 Buffer Operations (Data-Processing Path)

Section titled “1.2 Buffer Operations (Data-Processing Path)”
BenchmarkPayload SizeTimeThroughput
BufferWriteRead64 B4.42 μs27.6 MiB/s
BufferWriteRead256 B5.34 μs91.5 MiB/s
BufferWriteRead1200 B4.43 μs516.7 MiB/s
BufferWriteRead4096 B4.62 μs1.65 GiB/s
BufferWriteRead16384 B19.3 μs1.58 GiB/s
BufferEncodeVarInt—132 ns22.7M items/s

Analysis: Buffer operations excel at the typical QUIC packet size (1200B), with throughput above 500 MiB/s.

BenchmarkTimeThroughput
AckFrameEncode4.33 μs231K frames/s
AckFrameDecode4.54 μs220K frames/s
StreamFrameEncode9.05 μs110K frames/s

Analysis: ACK frame encode/decode at ~4.4μs fully satisfies QUIC protocol performance requirements.

BenchmarkTimeThroughput
QpackEncode (9 headers)6.19 μs1.46M headers/s
QpackDecode (5 headers)4.82 μs1.04M headers/s
HuffmanEncode (15 chars)139 ns103 MiB/s
HuffmanDecode (11 bytes)36.1 ns317 MiB/s

Analysis: QPACK encoding performs well. Huffman decoding is ~3× faster than encoding, indicating a highly efficient decode lookup table.

AllocatorSizeTimeSpeedup vs malloc
Poolallocator16 B2.10 ns6.2x
Poolallocator64 B2.10 ns6.0x
Poolallocator128 B2.06 ns6.1x
Poolallocator256 B2.06 ns12.6x
BlockMemoryPool1024 B7.13 ns1.5x
BlockMemoryPool4096 B7.13 ns1.5x
BlockMemoryPool16384 B6.99 ns1.6x
std::malloc16 B13.0 ns1.0x
std::malloc256 B28.5 ns1.0x
std::malloc4096 B10.6 ns1.0x

Key findings:

  • Poolallocator small objects (≤256B): 6–13× faster than malloc, at only ~2ns
  • BlockMemoryPool large blocks (1K–16K): 1.5× faster than malloc, at ~7ns
  • The custom allocators carry significant advantages for QUIC packet-processing memory management
BenchmarkTimeThroughput
PacketProcessingSimulation0.168 μs6.67 GiB/s

Analysis: The full per-packet processing path (allocate + write + parse header + read frame data) takes only 168 ns, theoretically supporting 5.95M packets/second.


Buffer CapacityCreation TimeChunk Count
1 KiB5.34 μs1
4 KiB5.37 μs1
16 KiB5.27 μs1
64 KiB23.8 μsmultiple
Block SizePool Size After 100 AllocationsPool Size After ReleaseAfter ReleaseHalf
1 KiB282814
4 KiB282814
16 KiB282814

Analysis: BlockMemoryPool’s batch allocation strategy is efficient. ReleaseHalf reclaims roughly 50% of idle memory.

MetricValue
Initial RSS~108 MB
RSS after 10K alloc/free cycles~108 MB
RSS growth0 KB

Conclusion: No memory growth observed across repeated alloc/free cycles, indicating the memory pools don’t leak.

Data WrittenChunk CountWrite Throughput
4 KiB12.47 GiB/s
16 KiB46.62 GiB/s
64 KiB1612.06 GiB/s
256 KiB6414.64 GiB/s

Analysis: Buffer Chain write throughput rises with data volume (amortizing object-creation overhead), reaching 14.6 GiB/s at 64 chunks.


WorkloadPoolallocatorstd::mallocSpeedup
Mixed alloc/free356 ns1108 ns3.1x
Packet CountTotal TimePer Packet
101.59 μs159 ns
10015.9 μs159 ns
1000159 μs159 ns

Analysis: The per-packet Buffer allocate + use + free overhead holds steady at 159 ns, scaling linearly.

Initial CapacityRequested BlocksFinal Pool SizePool Size After Release
4 blocks102 (idle)12 (returned to pool)
4 blocks50212
4 blocks200020
4 blocks500020

The following thresholds serve as reference criteria for performance-regression detection (comparing deviation between two runs’ JSON reports):

Metric CategoryThresholdNote
Default threshold15%beyond this, flag as regression
Critical paths10%packet processing, frame encode/decode
Allocators20%memory-allocation operations (more sensitive to external factors)

Optimization directions based on the baseline data:

  1. Buffer creation overhead (~5μs): consider a Buffer object pool to avoid creating a new Buffer per packet
  2. StreamFrame encoding (9μs): 2× slower than AckFrame (4.3μs); there may be room for optimization
  3. QPACK encoding (6μs/request): for high-frequency HTTP/3 requests, the dynamic-table hit rate is key
  4. BlockMemoryPool thread safety: lock contention under multi-threading; consider a per-thread pool

终端窗口
# Build
cmake -B build -DENABLE_PERF_TESTS=ON -DCMAKE_BUILD_TYPE=RelWithDebInfo
cmake --build build -j
# Run the CPU hotspot benchmarks
./build/bin/perf/cpu_hotspot_test
# Run the memory benchmarks
./build/bin/perf/memory_baseline_test
# Run the memory-pool efficiency analysis
./build/bin/perf/memory_pool_efficiency_test
# Output JSON reports
./build/bin/perf/cpu_hotspot_test --benchmark_format=json --benchmark_out=perf_results/cpu_hotspot.json
# Save a JSON report as the performance baseline (compare future runs against it)
./build/bin/perf/cpu_hotspot_test --benchmark_format=json --benchmark_out=baseline/cpu_hotspot.json
# Sample-profile for 30 seconds and generate a flame graph (collapsed stacks feed
# flamegraph.pl directly)
./build/bin/perf/profile_decode_packets --seconds 30 --out /tmp/decode_stacks.raw
python3 test/perf/tools/resolve_stacks.py /tmp/decode_stacks.raw
# ASan memory analysis (build separately with -DSANITIZER=asan, then run unit tests)
cmake -B build-asan -DSANITIZER=asan -DCMAKE_BUILD_TYPE=RelWithDebInfo
cmake --build build-asan -j
./build-asan/bin/quicx_utest

OptionDefaultDescription
ENABLE_PERF_TESTSONbuild the performance tests (test/perf, including the sampling profiler tools)
SANITIZER(empty)one of asan / ubsan / tsan; enables the corresponding sanitizer build

Note: perf targets carry their own profiling-friendly compile flags (-O2 -g -fno-omit-frame-pointer); no separate profiling switch is needed.