Metrics Monitoring System
quicX Metrics provides comprehensive observability for the quicX QUIC/HTTP3 implementation: 54+ metrics covering the UDP, QUIC, and HTTP/3 layers, lock-free design, Prometheus format export. Every metric operation is O(1) — atomics avoid lock contention, pre-allocated slots mean zero heap allocation, a single update costs < 10ns; core metrics are auto-instrumented and can be toggled at runtime. This document attempts to answer the following questions:
- How does the metrics system stay zero-overhead? — see “System Architecture” and “Performance Guarantees”;
- How are the 54+ metrics organized, and what does each cover? — see the 13 functional categories in “Metric Categories”;
- How to integrate and export in a project? — see “Usage Guide”.
1. System Architecture
Section titled “1. System Architecture”Core Components
Section titled “Core Components”┌─────────────────────────────────────────────────┐│ Application Code ││ (UDP, QUIC, HTTP/3, Memory Pool, etc.) │└────────────────┬────────────────────────────────┘ │ Metrics::CounterInc() │ Metrics::GaugeSet() ▼┌─────────────────────────────────────────────────┐│ Metrics Registry (Lock-Free) ││ ┌──────────────┐ ┌──────────────┐ ││ │ Counter │ │ Gauge │ ││ │ (Atomic) │ │ (Atomic) │ ││ └──────────────┘ └──────────────┘ │└────────────────┬────────────────────────────────┘ │ ExportPrometheus() ▼┌─────────────────────────────────────────────────┐│ Prometheus Exporter ││ # TYPE metric_name counter ││ metric_name{labels} value │└─────────────────────────────────────────────────┘Implementation Principles
Section titled “Implementation Principles”1. Lock-Free Slot Allocation
Section titled “1. Lock-Free Slot Allocation”// Pre-allocated slot arraystd::vector<MetricSlot> slots_; // Allocated at initialization
// O(1) registrationMetricID RegisterCounter(name, help) { size_t id = next_id_.fetch_add(1); // Atomic increment slots_[id] = MetricSlot{name, help, COUNTER}; return id;}2. Atomic Operation Updates
Section titled “2. Atomic Operation Updates”// Counter increment - lock-freevoid CounterInc(MetricID id, uint64_t delta = 1) { slots_[id].value.fetch_add(delta, std::memory_order_relaxed);}
// Gauge set - lock-freevoid GaugeSet(MetricID id, uint64_t value) { slots_[id].value.store(value, std::memory_order_relaxed);}3. Efficient Export
Section titled “3. Efficient Export”std::string ExportPrometheus() { std::ostringstream oss; for (auto& slot : slots_) { if (slot.type == COUNTER) { oss << "# TYPE " << slot.name << " counter\n"; oss << slot.name << " " << slot.value.load(std::memory_order_relaxed) << "\n"; } // ... Gauge, Histogram } return oss.str();}2. Metric Categories
Section titled “2. Metric Categories”1. UDP Layer Metrics (6 metrics)
Section titled “1. UDP Layer Metrics (6 metrics)”| Metric Name | Type | Description |
|---|---|---|
udp_packets_rx | Counter | Total UDP packets received |
udp_packets_tx | Counter | Total UDP packets sent |
udp_bytes_rx | Counter | Total UDP bytes received |
udp_bytes_tx | Counter | Total UDP bytes sent |
udp_dropped_packets | Counter | Total UDP packets dropped |
udp_send_errors | Counter | Total UDP send errors |
Purpose: Monitor network layer health, identify network congestion and packet loss issues.
2. QUIC Connection Layer Metrics (5 metrics)
Section titled “2. QUIC Connection Layer Metrics (5 metrics)”| Metric Name | Type | Description |
|---|---|---|
quic_connections_active | Gauge | Current active connections |
quic_connections_total | Counter | Total connections created |
quic_connections_closed | Counter | Total connections closed |
quic_handshake_success | Counter | Successful handshakes |
quic_handshake_fail | Counter | Failed handshakes |
Purpose: Monitor connection lifecycle, evaluate handshake success rate.
3. QUIC Packet Layer Metrics (6 metrics)
Section titled “3. QUIC Packet Layer Metrics (6 metrics)”| Metric Name | Type | Description |
|---|---|---|
quic_packets_rx | Counter | Total QUIC packets received |
quic_packets_tx | Counter | Total QUIC packets sent |
quic_packets_retransmit | Counter | Total retransmitted packets |
quic_packets_lost | Counter | Total packets lost |
quic_packets_dropped | Counter | Total dropped packets |
quic_packets_acked | Counter | Total packets acknowledged |
Purpose: Monitor transport layer reliability, calculate packet loss and retransmission rates.
4. QUIC Stream Layer Metrics (7 metrics)
Section titled “4. QUIC Stream Layer Metrics (7 metrics)”| Metric Name | Type | Description |
|---|---|---|
quic_streams_active | Gauge | Current active streams |
quic_streams_created | Counter | Total streams created |
quic_streams_closed | Counter | Total streams closed |
quic_streams_bytes_rx | Counter | Total stream bytes received |
quic_streams_bytes_tx | Counter | Total stream bytes sent |
quic_streams_reset_rx | Counter | RESET frames received |
quic_streams_reset_tx | Counter | RESET frames sent |
Purpose: Monitor stream management and data transmission, identify stream anomalies.
5. RTT Performance Metrics (3 metrics)
Section titled “5. RTT Performance Metrics (3 metrics)”| Metric Name | Type | Description |
|---|---|---|
rtt_smoothed_us | Gauge | Smoothed RTT (microseconds) |
rtt_variance_us | Gauge | RTT variance (microseconds) |
rtt_min_us | Gauge | Minimum RTT (microseconds) |
Purpose: Monitor network latency, evaluate connection quality.
6. Congestion Control Metrics (6 metrics)
Section titled “6. Congestion Control Metrics (6 metrics)”| Metric Name | Type | Description |
|---|---|---|
congestion_window_bytes | Gauge | Current congestion window (bytes) |
congestion_events_total | Counter | Total congestion events |
slow_start_exits | Counter | Slow start exits |
bytes_in_flight | Gauge | Bytes in flight |
pacing_rate_bytes_per_sec | Gauge | Pacing rate (bytes/second) |
pacing_delay_us | Histogram | Pacing delay (microseconds) |
Purpose: Monitor congestion control algorithm, optimize throughput.
7. Error Statistics Metrics (4 metrics)
Section titled “7. Error Statistics Metrics (4 metrics)”| Metric Name | Type | Description |
|---|---|---|
errors_protocol | Counter | Protocol errors |
errors_internal | Counter | Internal errors |
errors_flow_control | Counter | Flow control errors |
errors_stream_limit | Counter | Stream limit errors |
Purpose: Monitor system health, quickly identify issues.
8. Flow Control Metrics (2 metrics)
Section titled “8. Flow Control Metrics (2 metrics)”| Metric Name | Type | Description |
|---|---|---|
quic_flow_control_blocked | Counter | Connection-level flow control blocks |
quic_stream_data_blocked | Counter | Stream-level flow control blocks |
Purpose: Monitor flow control state, optimize window sizes.
9. Timeout Metrics (2 metrics)
Section titled “9. Timeout Metrics (2 metrics)”| Metric Name | Type | Description |
|---|---|---|
idle_timeout_total | Counter | Idle timeouts |
pto_count_total | Counter | PTO timeouts |
Purpose: Monitor timeout events, adjust timeout parameters.
10. HTTP/3 Metrics (4 metrics)
Section titled “10. HTTP/3 Metrics (4 metrics)”| Metric Name | Type | Description |
|---|---|---|
http3_requests_total | Counter | Total HTTP/3 requests |
http3_requests_active | Gauge | Current active requests |
http3_requests_failed | Counter | Failed requests |
http3_push_promises_rx | Counter | Push promises received |
Purpose: Monitor HTTP/3 business metrics, evaluate service quality.
11. Memory Pool Metrics (4 metrics)
Section titled “11. Memory Pool Metrics (4 metrics)”| Metric Name | Type | Description |
|---|---|---|
mem_pool_allocated_blocks | Gauge | Allocated blocks |
mem_pool_free_blocks | Gauge | Free blocks |
mem_pool_allocations | Counter | Allocation count |
mem_pool_deallocations | Counter | Deallocation count |
Purpose: Monitor memory usage, optimize memory pool configuration.
12. Frame Statistics Metrics (2 metrics)
Section titled “12. Frame Statistics Metrics (2 metrics)”| Metric Name | Type | Description |
|---|---|---|
frames_rx_total | Counter | Total frames received |
frames_tx_total | Counter | Total frames sent |
Purpose: Monitor protocol layer activity, analyze communication patterns.
13. ACK Related Metrics (3 metrics)
Section titled “13. ACK Related Metrics (3 metrics)”| Metric Name | Type | Description |
|---|---|---|
ack_delay_us | Histogram | ACK delay (microseconds) |
ack_ranges_per_frame | Histogram | ACK ranges per frame |
ack_frequency | Gauge | ACK frequency (ACKs per second) |
Purpose: Monitor ACK behavior, optimize acknowledgment strategy.
3. Usage Guide
Section titled “3. Usage Guide”Initialization
Section titled “Initialization”#include <quicx/common/metrics.h>
// 1. Configure Metricsquicx::MetricsConfig config;config.enable_ = true; // Enable metricsconfig.initial_slots_ = 1024; // Initial slot countconfig.prefix_ = "quicx_"; // Metric name prefix
// 2. Initialize Metrics systemquicx::Metrics::Initialize(config);Export Prometheus Format
Section titled “Export Prometheus Format”// Get metrics data in Prometheus formatstd::string metrics_data = quicx::Metrics::ExportPrometheus();
// Write to filestd::ofstream file("/var/lib/prometheus/quicx.prom");file << metrics_data;file.close();
// Or serve via HTTP endpoint// (See HTTP/3 Metrics Endpoint section)HTTP/3 Metrics Endpoint
Section titled “HTTP/3 Metrics Endpoint”#include <quicx/http3/if_server.h>
// Create HTTP/3 serverauto server = quicx::IServer::Create(settings);
// Configure serverquicx::Http3ServerConfig config;config.quic_config_.cert_file_ = "server.crt";config.quic_config_.key_file_ = "server.key";
// Enable metrics endpointconfig.metrics_.http_enable_ = true;config.metrics_.http_path_ = "/metrics";
// Initialize and startserver->Init(config);server->Start("0.0.0.0", 8443);Access metrics:
# Using HTTP/3 clientcurl --http3 https://localhost:8443/metrics4. Performance Guarantees
Section titled “4. Performance Guarantees”Benchmark Results
Section titled “Benchmark Results”Benchmark Time CPU-------------------------------------------------CounterInc/1 8.2 ns 8.2 nsCounterInc/100 820 ns 820 nsGaugeSet/1 7.5 ns 7.5 nsGaugeSet/100 750 ns 750 nsExportPrometheus/100 45.2 µs 45.2 µsExportPrometheus/1000 452.0 µs 452.0 µsPerformance Characteristics
Section titled “Performance Characteristics”- Ultra-Low Latency: Single update < 10ns
- Linear Scaling: Performance scales linearly with metric count
- Zero Contention: Lock-free design, no contention in multi-threaded scenarios
- Memory Efficient: Pre-allocated, no runtime allocation
Memory Footprint
Section titled “Memory Footprint”Per metric slot: ~128 bytes1000 metrics: ~128 KBExport buffer: ~100 KB (temporary)5. Best Practices
Section titled “5. Best Practices”1. Configure Slot Count Appropriately
Section titled “1. Configure Slot Count Appropriately”// Configure based on expected metric countconfig.initial_slots_ = expected_metrics * 1.5; // Leave 50% headroom2. Export Regularly
Section titled “2. Export Regularly”// Export every 15 seconds (Prometheus default scrape interval)std::thread exporter([]{ while (running) { std::string data = Metrics::ExportPrometheus(); WriteToFile("/var/lib/prometheus/quicx.prom", data); std::this_thread::sleep_for(std::chrono::seconds(15)); }});3. Monitor Key Metrics
Section titled “3. Monitor Key Metrics”Priority monitoring:
- Connection count (
quic_connections_active) - Packet loss rate (
quic_packets_lost / quic_packets_tx) - RTT (
rtt_smoothed_us) - Error rate (
errors_*) - Throughput (
quic_streams_bytes_*)
4. Set Up Alerts
Section titled “4. Set Up Alerts”# Prometheus alert rules examplegroups: - name: quicx_alerts rules: - alert: HighPacketLoss expr: rate(quic_packets_lost[5m]) / rate(quic_packets_tx[5m]) > 0.05 annotations: summary: "High packet loss rate (> 5%)"
- alert: HighRTT expr: rtt_smoothed_us > 100000 # > 100ms annotations: summary: "High RTT detected"
- alert: TooManyErrors expr: sum(rate(errors_protocol[5m])) > 10 annotations: summary: "High error rate"6. Troubleshooting
Section titled “6. Troubleshooting”Issue: Metrics Not Updating
Section titled “Issue: Metrics Not Updating”Cause: Metrics not initialized or disabled
Solution:
// Ensure initializationMetrics::Initialize(config);
// Ensure enabledconfig.enable_ = true;Issue: Export Data Empty
Section titled “Issue: Export Data Empty”Cause: No metrics registered
Solution:
// Ensure InitializeStandardMetrics() was called// This is automatically called in Metrics::Initialize()Issue: Performance Degradation
Section titled “Issue: Performance Degradation”Cause: Insufficient slots, causing reallocation
Solution:
// Increase initial slot countconfig.initial_slots_ = 2048; // Or larger7. Extension Development
Section titled “7. Extension Development”Adding Custom Metrics
Section titled “Adding Custom Metrics”// 1. Declare in metrics_std.hstruct MetricsStd { static MetricID MyCustomMetric;};
// 2. Define in metrics_std.cppMetricID MetricsStd::MyCustomMetric = kInvalidMetricID;
// 3. Register in InitializeStandardMetrics()MetricsStd::MyCustomMetric = Metrics::RegisterCounter("my_custom_metric", "My custom metric");
// 4. Use in codeMetrics::CounterInc(MetricsStd::MyCustomMetric);Adding Histogram Support
Section titled “Adding Histogram Support”// Register HistogramMetricID latency_hist = Metrics::RegisterHistogram( "request_latency_us", "Request latency in microseconds", {10, 50, 100, 500, 1000, 5000} // buckets);
// Record observationMetrics::HistogramObserve(latency_hist, latency_value);