The Last Engineer Who Knew

Network firefighting is no longer a strategy.

Somewhere in your network right now there is a shelf, a card, or a ring segment that went end-of-life before some of your engineers graduated high school. And somewhere in your organisation — maybe — there is one person who still remembers how it works. How the protection switching actually behaves. Which tributaries are mapped where. Why that one node was configured differently from every other node in the ring, and what happens if you normalise it.

This is about what happens when that person leaves.

Why is this happening all at once?

Three forces are converging, and the damage compounds rather than adds up.

Retirement is removing the deepest layer. The engineers who designed and built the networks carrying today’s traffic are reaching the end of their careers, and what they carry lives in their heads rather than in a documentation system. Downsizing is cutting into the layer beneath them, and the arithmetic is particularly cruel: the most expensive employees are the ones with the longest tenure and the deepest context, which makes them the first names on a reduction list. Expert flight removes the third layer — the mid-career engineers who might have stepped up are being recruited by cloud providers, financial services, and the AI sector, all of which pay a premium for exactly the skills telecom needs.

Each vector accelerates the other two. Retirements remove the knowledge. Downsizing removes the experienced replacements. Expert flight removes the ambitious mid-career engineers who would have absorbed it. What remains is a workforce that is simultaneously thinner, less experienced, and facing more complexity than the one it replaced.

Why can’t you hire or train your way out of it?

Because no operator runs one technology. A working network today spans four to six generations at once — TDM, SONET, MPLS, GPON, XGS-PON, software-defined transport, 4G, 5G — and each generation carries its own CLI syntax, alarm semantics, provisioning workflows, and failure modes. Each vendor inside each generation adds another layer: a Calix AXOS OLT does not behave like a Nokia ISAM, and a Cisco IOS-XR router is a fundamentally different operational experience from an IOS-XE device despite the shared brand.

The expertise required spans hundreds of vendor, device type, and generation combinations. No single engineer masters all of it. The ones who come closest — who can contextualise a new alarm because they worked the older platform it inherited its behaviour from — are precisely the ones retiring, being downsized, or leaving.

What does the knowledge gap actually cost?

The Uptime Institute, drawing on twenty-five years of longitudinal data, attributes between 66% and 80% of downtime incidents to human error directly or indirectly. The detail that matters is what kind. Among organisations that suffered a major outage caused by human error, 85% traced the root cause to staff failing to follow procedures — or to the procedures themselves being flawed, outdated, or nonexistent.

That proportion is not stable. Uptime’s data shows outages caused by procedure failures rising from 48% to 58% in a single year. And 80% of operators surveyed said their most recent serious outage could have been prevented with better management, process, or configuration.

This is not a carelessness problem. It is a “the procedure does not exist, is outdated, or cannot be found at 3am” problem — which is to say it is a knowledge management problem, and it is getting worse precisely because the people who carried those procedures in their heads are the ones leaving.

Why did SharePoint, Confluence, and the runbook binder fail?

Operators have tried this before. Wikis, document portals, vendor libraries, binder-based SOPs. Every one failed for the same structural reason: they required people under operational pressure to search correctly, to maintain content proactively, and to organise information consistently. Under the stress of a 3am outage, none of those behaviours happen.

An engineer spending a third of a shift hunting for a runbook is an engineer not resolving an outage. A search that fails at 3am does not merely waste time — it forces the engineer to improvise, escalate, or guess. That is how the 85% procedure-failure figure gets generated.

What actually changed?

Retrieval-augmented generation is a structural break rather than an incremental improvement, and the difference is specific.

Previous tools required the engineer to construct a precise keyword search against an organised repository. Wrong terminology, wrong filename, wrong folder — nothing useful comes back. A retrieval-grounded system works from intent. An engineer who types “the ONT on Elm Street keeps dropping” gets the right troubleshooting procedure without knowing the alarm name, the model number, or which document holds the answer.

Critically, the documentation does not have to be clean first. No six-month migration project before the system delivers value — which is exactly the requirement that killed its predecessors.

What separates a useful system from another unused tool?

Three capabilities, and a system missing any one of them becomes shelfware.

Capture what exists. Ingest the runbooks, SOPs, MOPs, vendor manuals, and ticket history as they are, and make them queryable immediately. Identify what is missing. Analyse documentation coverage against the actual device inventory — which device types have no procedures, which alarm categories have no runbooks, which queries consistently return nothing. That turns knowledge management from a passive repository into an investment plan. Document as you go. When an engineer resolves something that required tribal knowledge, capture the resolution as a draft procedure for a shift lead to approve. Every resolved incident becomes an asset, and the knowledge base grows with operational tempo instead of depending on documentation sprints that never get prioritised.

This is the design behind Wade, the institutional-memory agent in Rapax. Wade holds the operator’s procedures and returns a confidence score with every answer — which is what lets the other agents act on retrieved knowledge without guessing, and what stops the system inventing a procedure when none exists.

The timing argument is simple. The engineers carrying your institutional knowledge are still here today. In three to five years many of them will not be, and capturing knowledge now costs a fraction of reconstructing it later — assuming reconstruction is possible at all.

Get the white paper

The full paper adds the sourced industry data behind the figures above, the six-phase implementation roadmap with success metrics at each stage, and direct answers to the three objections this always raises — we have tried knowledge management before, can we trust AI answers for critical procedures, and what happens to our proprietary data.

ContactUs TelecomKnowledgeCrisisKB

About Rapax

Rapax is an AI-native network service assurance and automation platform for telecom operators, bringing fault, performance, topology, and service management into one system worked by six AI agents. Rapax is a business unit of Citus Technologies, LLC.

Shawn Ennis is the Founder & CEO of Citus Technologies and the founder of Rapax. He started his career building SONET rings, spent 25 years in telecom operations, holds 12 patents in network management and service assurance, and previously founded Assure1 — acquired by Oracle in 2021. He co-hosts the Transformation Leaders Podcast.

Not ready to download? Book fifteen minutes — no prep, no deck. cal.com/shawn-ennis · sales@rapax.app · More white papers