Lightweight Network SLA Monitoring Platform (LNMP)
Building LNMP v2.0: Sugiyama DAG Topology, TimescaleDB Compression, and Session Security
How production multi-user deployments led to LNMP v2.0: implementing a 4-phase Sugiyama topology engine, TimescaleDB compression, top-of-minute write semaphores, and solving in-memory authentication deadlocks.
In Part 1 and Part 2 of this series, we developed the Lightweight Network Monitoring Platform (LNMP) and refined its failure isolation capabilities through adaptive Z-score baselines and differential root cause analysis. While Versions 1.0 and 1.5 were validated primarily inside virtual simulation environments (GNS3 network topologies), moving LNMP into real-world, multi-user production infrastructure exposed an entirely new set of scaling constraints: force-directed graph wire crossings, database connection pool exhaustion at top-of-the-minute polling boundaries, multi-gigabyte time-series growth, and in-memory authentication deadlocks. LNMP Version 2.0 addresses these operational challenges with a 4-phase Sugiyama hierarchical topology engine, TimescaleDB chunk compression, write semaphores, and an overhauled session governance model.
LNMP v2.0 Architecture Overview
FastAPI Service API (v2.0)
IP-Scoped Lockouts • Sliding 2h Inactivity Sessions • Global Latency Middleware
Monitoring Engine (The Poller)
Write Semaphore (15) • Top-of-Minute Sync • Connection Pool Scaling (20/30)
4-Phase Sugiyama Graph Engine
BFS Longest-Path Layering • Barycenter Crossing Reduction • Gansner Alignment
From Lab Simulation to Production Reality
The shift from Version 1.5 to Version 2.0 represents the difference between running tests in a simulated lab and supporting concurrent human operators across live enterprise networks:
- Simulation vs Production Failure Modes: In GNS3 lab testing, we monitored clean synthetic topologies with predictable route tables and single-operator access. In production, real hardware generated unpredicted edge crossings on canvas maps, endpoints attempted database writes at the exact same microsecond, and multiple engineers accessing the dashboard concurrently exposed session deadlocks.
- Observability Deficits: When strange behavior occurred in production, our console-only logging lacked the granular request traces and latency metrics required to isolate slow database queries or authentication rejections.
- Storage Accumulation: Sustained 1-minute telemetry across hundreds of nodes rapidly consumed disk space, requiring native database compression rather than standard relational table storage.
Here is how the architectural baseline evolved from v1.5 to v2.0:
| Architectural Area | LNMP v1.5 (Baseline) | LNMP v2.0 (Current Architecture) |
|---|---|---|
| Topology Visualization | Physics-stabilized force-directed graph (frozen after 200 iterations). | 4-Phase Sugiyama Hierarchical Engine: BFS Longest-Path Layering + Barycenter Crossing Reduction + Gansner Coordinate Alignment. |
| Canvas Orientation | Fixed single-direction canvas. | Dynamic Layout Switcher: Instant animated toggle between Horizontal (Left-to-Right LR) and Vertical (Top-to-Bottom UD) views. |
| Time-Series Storage | Uncompressed TimescaleDB hypertables. | 7-Day Native Columnar Compression: Reduces disk footprint by 90%+ while keeping historical telemetry 100% queryable. |
| Baseline Automation | Manual continuous aggregate refresh queries. | Continuous Aggregate Policies: Scheduled hourly background refresh with crash-recovery catch-up. |
| Database Concurrency | Direct asynchronous writes on minute boundaries. | Top-of-Minute Write Semaphore (asyncio.Semaphore(15)): Eliminates connection pool exhaustion during polling bursts. |
| Authentication State | Global username-scoped in-memory lockout dictionaries. | IP-Scoped Lockouts (f"{client_ip}:{username}"): Isolates brute-force attacks without locking out legitimate admins on other IPs. |
| Session Lifecycle | Fixed token expiry with static active session rejection. | Sliding 2-Hour Inactivity Window + FIFO Token Eviction: Slides active tokens on requests; evicts oldest session when device quota (2) is met. |
| Out-of-Band Recovery | Required server restart to clear in-memory state. | Dedicated CLI Password Reset Tool: Out-of-band script (deploy/reset-admin-password.sh) clearing lockout states directly. |
| Credential Management | Custom inputs without autofill support. | Native Browser Autofill Architecture: Standard unnested HTML inputs with autocomplete and password visibility toggles. |
| Observability & Logging | Unbuffered stdout console output. | 150MB Auto-Rotating Logging Suite: Dual console and bounded file outputs (api.log, engine.log, error.log). |
| Lifecycle Scripts | Basic install/upgrade scripts. | Refined Upgrade & Decommission Pipelines: Systemd auto-start enforcement, smart config migrations, and safe uninstaller with SQL dump. |
4-Phase Sugiyama Topology Engine: Eliminating Wire Crossings
Why Force-Directed Graphs Broke in Production
In Version 1.5, we stabilized the topology canvas by freezing spring physics after 200 iterations. But once deployed against complex enterprise transit paths, force-directed layouts revealed severe visualization flaws:
- Hop Depth Inversion: Multi-hop paths drifted vertically, frequently placing Level 3 branch one the same level on the screen with Level 2 core switches.
- Diagonal Wire Crossings: Edges sliced across unrelated intermediate nodes, making it difficult for operators to trace actual link dependencies.
- Parent Misalignment: Core aggregation routers did not align geometrically above their dependent child clusters.
We evaluated several academic graph layout algorithms to fix this while preserving our existing in-memory DAG engine. I eventually opted for a based on Sugiyama et al. (1981) combined with the coordinate assignment heuristics of Gansner et al. (1993).
4-Phase Sugiyama & Gansner Topology Pipeline
BFS DAG Layering
Level(v) = max(L(u)+1)
Assigns every node to its exact discrete physical hop tier, consolidating shared gateways.
Crossing Reduction
edgeMinimization: true
Sugiyama barycenter heuristic sorts sibling nodes to ensure parallel, non-overlapping channels.
Gansner Alignment
blockShifting: true
Centers parent routers directly over child clusters and maintains spatial corridors between subtrees.
Spline Routing
cubicBezier Tangents
Channels smooth splines aligned dynamically with the active layout orientation (LR vs UD).
Phase 1: BFS DAG Longest-Path Layering
The engine traverses the in-memory DAG starting from the local monitoring gateway (Root, Level 0). Every downstream vertex receives an integer layer rank based on its maximum parent distance:
This mathematical constraint guarantees that every router, switch, and host is placed on its exact discrete physical hop tier, preventing backward or diagonal cross-tier jumps.
Phase 2: Barycenter Crossing Minimization
To minimize edge crossings between adjacent layers and , the engine orders vertices within each layer using the Sugiyama barycenter heuristic:
Vertices on layer are sorted by their computed barycenter values, ensuring that edges run in parallel downward channels without crossing.
Phase 3 & 4: Gansner Coordinate Alignment and Spline Routing
Using Gansner coordinate heuristics (blockShifting: true, parentCentralization: true), parent routers are positioned directly above the geometric center of their dependent child clusters. Spatial corridors are maintained between separate branch subtrees to prevent visual collisions.
Finally, edges are rendered using directional cubic Bézier splines with dynamic tangent constraints:
// vis-network hierarchical configuration in LNMP v2.0
const topologyOptions = {
layout: {
hierarchical: {
enabled: true,
direction: currentOrientation.value, // "UD" (Top-to-Bottom) or "LR" (Left-to-Right)
sortMethod: "directed",
nodeSpacing: 180,
levelSeparation: 160,
blockShifting: true,
edgeMinimization: true,
parentCentralization: true,
},
},
physics: {
hierarchicalRepulsion: {
centralGravity: 0.0,
springLength: 100,
nodeDistance: 150,
damping: 0.09,
},
},
edges: {
smooth: {
type: "cubicBezier",
forceDirection: currentOrientation.value === "UD" ? "vertical" : "horizontal",
roundness: 0.5,
},
},
};
Dynamic Horizontal (LR) ⇄ Vertical (UD) Switcher
Widescreen displays often benefit from left-to-right signal flow rather than tall top-to-bottom trees. LNMP v2.0 adds an instantaneous layout switcher as requested by a user that flips direction between UD and LR while adjusting spline tangent constraints dynamically, providing smooth visual transitions without page reloads.
Concurrency Protection & Top-of-Minute Write Semaphore
The Thundering Herd Problem in Production
In Version 1.5, polling cycles were synchronized to top-of-the-minute boundaries (:00 seconds). In lab tests with 10 endpoints, database writes were instantaneous. But when the system was tested the server against 200 endpoints, all 200 polling tasks finished their 10-ping sub-cycles at second 59 and attempted to write their summaries to PostgreSQL at the exact same microsecond. Luckly this was detected while stesstesting the machine in a testing environment.
This triggered severe connection pool exhaustion errors:
sqlalchemy.exc.TimeoutError: QueuePool limit of size 20 overflow 10 reached, connection timed out
The Semaphore & Pool Scaling Solution
LNMP v2.0 introduces a database write semaphore (asyncio.Semaphore(15)) inside the monitoring daemon to throttle concurrent write operations:
import asyncio
from sqlalchemy.ext.asyncio import AsyncSession
db_write_semaphore = asyncio.Semaphore(15)
async def commit_subcycle_summary(endpoint_id: str, summary_data: dict, session_factory):
"""Write subcycle telemetry with concurrency throttling."""
async with db_write_semaphore:
async with session_factory() as session:
async with session.begin():
await write_endpoint_event(session, endpoint_id, summary_data)
Combined with expanded pool parameters (pool_size=20, max_overflow=30) in database.py, this semaphore smooths top-of-minute write spikes into a controlled 200 ms execution window, eliminating connection pool timeouts.
Telemetry Datastore Optimization: TimescaleDB 7-Day Compression
Solving Multi-Year Hypertable Growth
At 1-minute polling intervals with 10-ping sub-cycles, monitoring 200 endpoints generates 288,000 telemetry rows daily (). In uncompressed relational tables, index maintenance and disk I/O degrade query speeds over multi-year retention windows.
LNMP v2.0 implements native TimescaleDB Columnar Hypertable Compression via migration 0005_v2_0_timescale_stability.py:
-- Configure columnar chunk compression
ALTER TABLE endpoint_events SET (
timescaledb.compress,
timescaledb.compress_segmentby = 'endpoint_id',
timescaledb.compress_orderby = 'start_time DESC'
);
-- Activate automated 7-day compression policy
SELECT add_compression_policy('endpoint_events', INTERVAL '7 days', if_not_exists => true);
Operational Impact
- 90%+ Storage Reduction: Historical metric chunks older than 7 days are compressed into columnar formats, reducing disk storage by over .
- Query Transparency: Compressed data remains fully queryable via standard PostgreSQL
SELECTqueries without manual decompression or application-level changes. - Automated Continuous Aggregate Policies: Historical baseline calculations are automated via background policies that handle hourly refreshes and recover automatically after server reboots, as outlined in the TimescaleDB Continuous Aggregate documentation:
SELECT add_continuous_aggregate_policy(
'node_historical_baselines',
start_offset => INTERVAL '30 days',
end_offset => INTERVAL '1 hour',
schedule_interval => INTERVAL '1 hour',
if_not_exists => true
);
Root Cause Analysis: Fixing Multi-User Auth Deadlocks and Lockouts
The Production Incident
During multi-user onboarding and production testing, a frustrating bug surfaced: operator accounts were suddenly unable to log in, receiving 403 Forbidden or 401 Unauthorized errors despite entering valid credentials. The issue persisted until I manually executed systemctl restart netmon-api.
A deep root cause investigation revealed three distinct architectural flaws in the v1.5 authentication subsystem:
| Root Cause Flaw | Operational Defect & Resolution Implemented |
|---|---|
| Global Unscoped Lockout State | Failed attempts were stored in Python RAM keyed solely by username (_failed_attempts[username]). A typo on one device locked out the account for all operators globally. Fixed by scoping lockouts to f"{client_ip}:{username}". |
| In-Memory Session Quota Deadlock | Closing browser tabs without clicking “Sign Out” left tokens in _active_user_sessions. When the device limit was reached, new logins from other laptops were rejected. Fixed with sliding 2h window + FIFO token eviction. |
| Lack of Out-of-Band Reset Tooling | Admins locked out of the web UI had no CLI tool to reset passwords without killing the Python process. Fixed with standalone deploy/reset-admin-password.sh. |
1. IP-Scoped Failed Login Protection
In v1.5, the lockout tracker was stored globally by username:
# Old Legacy Vulnerable Pattern (v1.5)
_failed_attempts[username] = {"count": 5, "locked_until": ...}
If an operator left a background browser tab open with stale credentials, or if a colleague mistyped the administrator password five times, the admin account was locked globally for all operators across the enterprise. The only way to clear _failed_attempts was restarting the Python daemon.
LNMP v2.0 scopes lockout tracking to the client origin IP:
# Upgraded Production Pattern (v2.0)
def record_failed_login(client_ip: str, username: str) -> bool:
"""Track failed attempts per IP:username pair to prevent global lockouts."""
key = f"{client_ip}:{username}"
attempts = login_tracker.get(key, 0) + 1
login_tracker[key] = attempts
if attempts >= 5:
lockout_manager.lock(key, duration_minutes=15)
logger.warning(f"Security Alert: IP-scoped lockout triggered for {key}")
return True
return False
If an attacker or a single misconfigured host fails five logins, only that specific IP address is blocked for 15 minutes; administrators connecting from other subnets remain completely unaffected.
2. Sliding Inactivity Window & FIFO Session Eviction
In v1.5, active sessions were tracked in an in-memory dictionary. If an engineer closed their laptop lid without explicitly clicking “Sign Out”, their token remained active in the registry until hard expiration. Once an account hit its concurrent session cap, logins from new workstations were rejected.
LNMP v2.0 implements OWASP Session Management Guidelines and RFC 7519 JWT standards:
- Sliding 2-Hour Inactivity Window: Every active HTTP request slides the cookie expiration timestamp forward. Sessions idle for longer than 120 minutes expire automatically.
- FIFO Token Eviction: Each login generates a unique session ID (
jti). When an account logs in on a third device, the system automatically invalidates the oldest session token rather than rejecting the new login.
3. Out-of-Band CLI Recovery Tool & Native Autofill
- CLI Password Reset Tool: Version 2.0 includes
deploy/reset-admin-password.sh, allowing administrators to reset credentials and clear lockout state directly from the server terminal without restarting services. - Native Browser Password Autofill: Refactored login and password change forms into standard, unnested HTML
<form>elements with explicitnameandautocompleteattributes (username,current-password,new-password), enabling instant one-click credential saving across password managers (Bitwarden, 1Password).
Observability, Logging & Lifecycle Tooling
150MB Auto-Rotating Logging Suite
In production, console-only logging proved insufficient for debugging transient deployment issues. LNMP v2.0 introduces structured, rotating log files in /var/log/netmon/ using Python’s RotatingFileHandler:
api.log: HTTP access logs, status codes, and execution latencies (max 50 MB, 2 backups).engine.log: Polling cycles, sub-cycle health summaries, and state machine commits (max 50 MB, 2 backups).error.log: Diagnostic warnings, stack traces, and 401/403 security alerts (max 50 MB, 2 backups).
The total logging footprint is strictly capped at 150 MB.
Global Latency Middleware
An asynchronous FastAPI middleware logs the client IP, HTTP method, path, response status, and processing time for every request:
@app.middleware("http")
async def track_latency_middleware(request: Request, call_next):
start_time = time.perf_counter()
response = await call_next(request)
duration_ms = (time.perf_counter() - start_time) * 1000
logger.info(
f"{request.client.host} - {request.method} {request.url.path} "
f"[{response.status_code}] ({duration_ms:.2f}ms)"
)
return response
Upgrade Pipeline and Decommission Utility
While the decommission utility was initially introduced in v1.5, production testing revealed that upgrade scripts failed to handle service restarts cleanly. In v2.0:
deploy/upgrade.sh: Automatically dumps a pre-upgrade database backup, migrates/etc/netmon/config.tomldefaults without overwriting secrets, runs Alembic migrations, and gracefully restarts systemd units (systemctl restart netmon-*).deploy/uninstall.sh: Prompts for explicit confirmation, generates a pre-removal PostgreSQL database dump in/var/backups/netmon/, stops systemd services, removes Nginx reverse proxy configurations, and cleans up application binaries without accidental data loss.
Zero-Downtime Upgrade Pipeline for v2.0
Upgrading an existing installation to v2.0 is fully automated through deploy/upgrade.sh:
# Execute automated zero-downtime upgrade
sudo ./deploy/upgrade.sh
Zero-Downtime Upgrade Pipeline (./deploy/upgrade.sh)
Pre-Upgrade Dump
Dumps live database to /var/backups/netmon/ before changes.
Config Migrator
Applies v2.0 defaults without overwriting secrets.
Code & Assets
Pulls git updates, syncs packages, and builds Vue bundle.
DB Migrations
Runs migration 0005 for compression & agg policies.
Restart Units
Restarts daemons, reloads Nginx, and verifies health.
Conclusion & Ongoing Refinements
LNMP Version 2.0 demonstrates how real-world multi-user production deployments expose bottlenecks that sterile lab simulations never reveal. By replacing force-directed graphs with 4-phase Sugiyama topology layering, implementing 7-day TimescaleDB columnar compression, protecting database writes with semaphores, and eliminating in-memory authentication deadlocks, the platform achieves enterprise stability with minimal resource overhead.
Ongoing engineering efforts are focused on extending unprivileged topology discovery across Linux VRF namespaces and refining heuristic filters for transit carrier ICMP rate limiting.
The complete open-source codebase for LNMP v2.0 (Beta), including the poller daemon, FastAPI backend, Alembic migrations, and Vue 3 frontend, is available on GitHub:
- Repository: https://github.com/dc8official/lnmp.git