Your Business Already Has a Brain. AI Just Gave You Access.

The on-call scenario

AI Incident Knowledge Base for MSPs (Technician working in a server room on their laptop)

It’s 22:00. A monitoring alert lands. Your senior engineer, the one who set up half the client environments, is on leave. The junior tech on shift has never seen this alert pattern. They have thirty minutes before the client notices the outage themselves.

There are two ways this goes.

Version one: junior tech opens Slack, pings the group, waits for someone to remember. Digs through last year’s tickets in the PSA. Finds a similar-looking incident, can’t tell if the fix applies. Escalates. Loses forty minutes to context assembly before anyone touches the actual problem.

Version two: junior tech opens the internal assistant, describes the alert. The assistant surfaces three past incidents that matched the pattern, the resolution notes for each, and a flag that says “this environment uses a non-standard firewall config, see runbook 44.” Tech applies the documented fix in five minutes.

The difference is not seniority. It’s whether your institutional memory is queryable.

The mistake most SA IT businesses make with AI

When IT people ask me how I use AI, they expect me to say “for writing scripts” or “for summarising Sentinel alerts.” Both are fine. Both are shallow.

Here’s what actually changes an MSP: I stopped asking AI for answers, and I started asking it about my own incident history.

Every alert. Every root cause. Every fix that worked. Every client environment quirk. Every “we tried this and it broke worse.” All of that already exists in your business. It sits in ticket comments, in a Slack thread from six months ago, in a Confluence page nobody updated, in a technician’s OneNote you can’t access because they left in April.

That gap between “we know this somewhere” and “we can act on it in the middle of a P1” is your biggest hidden cost. AI closes it.

What this looks like at MSP scale

I run a managed IT arm alongside a marketing agency. Between them: a portfolio of clients across financial services, e-commerce and industrial. Each with a different security stack, a different backup posture, a different set of quirks in the environment.

The only way that works with a small team is that we don’t carry the context in our heads. We carry a system that surfaces it.

Every environment discovery goes into one place. Every incident post-mortem. Every configuration change. Every “we noticed the client’s cyber-insurer changed underwriter this quarter.” That library is indexed by a retrieval layer. When someone opens a ticket for that client, the assistant sees the last six months of context before anyone types a diagnostic command.

That’s not a productivity hack. That’s a different kind of MSP.

Why "just use ChatGPT" doesn't get you here

Most IT teams’ AI use is the browser tab. Paste an error message, get a generic suggestion, close the tab. That interaction has amnesia baked in. It knows nothing about your client’s specific environment, their compliance posture, or that the exact fix ChatGPT is suggesting was tried in this environment last year and made things worse.

So the output is generic advice. Sometimes correct in the abstract. Wrong for this environment. In IT, generic-correct is still dangerous when applied to a specific stack.

The problem is not the AI. The problem is that you’re asking a stranger to touch someone else’s production.

The three-layer setup that changes this

Layer 1: One home for every piece of environment knowledge. Pick one place where every meaningful thing about a client environment lives as searchable text. Environment discoveries, incident notes, config changes, quirks. The requirement isn’t the tool. It’s that “documented somewhere” and “in that place” mean the same thing. Ticket comments in the PSA don’t count if nobody can search them across clients.

Layer 2: A capture discipline. After every incident, three lines get written down. What happened. What fixed it. What we’d do differently next time. After every environment change, one line: what changed and why. If those lines never get captured, the incident effectively didn’t happen for anyone who wasn’t there. This discipline is the whole battle.

Layer 3: A retrieval layer over the whole thing. This is where AI enters. Point an assistant at the library and let it answer environment-specific questions using only what’s in there. When on-call asks “have we seen this before in this client,” the answer arrives in seconds.

Once these three layers are in place, the change isn’t just faster tickets. The change is that a technician in year one can approach year-five capability, because the year-five knowledge is loaded into the environment they work in.

Three failure patterns that kill this

Trusting the PSA as your knowledge base. Ticket comments are transactional. They aren’t searchable at scale, they don’t cross-reference across clients, and they capture the fix but rarely the reasoning. The PSA is where work happens. The knowledge base is where lessons live. Different jobs.

Asking AI to organise the knowledge for you. AI can’t invent post-mortems your team never wrote. If your engineers close tickets without documenting root cause, no retrieval layer will magically produce it. Capture is on the team. Retrieval is where AI earns its keep.

Confusing the AI answer with the source. Before applying an AI-surfaced fix to production, click through to the original incident note. If the note is thin, treat the answer as a hint, not a plan. This one habit prevents most AI-related MSP mistakes.

What this does for cyber insurance and POPIA questions

The effect goes beyond the on-call desk. When a cyber insurer sends the renewal questionnaire, “do you have documented incident response procedures” stops being a yes-with-fingers-crossed and becomes a yes-with-evidence. When a POPIA compliance question lands, “how do we handle personal data in the client environment” has a documented answer per client, not a fresh assembly job.

Both of these are commercial. Cyber insurers price better and cover more when they see documented process. Compliance posture is increasingly a sales lever, not just a defensive one. A queryable knowledge base is the raw material for both.

The unfair advantage nobody talks about

Every MSP in your market has access to the same public AI. The public AI gives everyone the same generic answers. That’s not a moat. That’s a tie.

Your moat is the environment knowledge that only exists inside your business. Your specific clients, your specific fixes, your specific quirks. That knowledge is already yours. It’s already written down, badly, in six different systems.

The MSPs that compound the fastest over the next three years are likely to be the ones that take that pile seriously enough to make it queryable. Not because AI is magical. Because for the first time, the cost of turning years of scattered incident notes into a working on-call brain is close to zero.

Do the capture. Build the library. Point the retrieval layer at it. That’s the whole play.

Everything else is script generation, which is nice, and completely beside the point.

Want your incident library working for you at 22:00?

If this landed, the practical next step isn’t more monitoring. It’s turning the institutional knowledge your team already has into something the on-call desk can query. Gknect Managed IT builds this into every engagement, because the runbook is the moat.

Reach out via gknect.com.

Part 1 of a 4-part series on how Marcos Diez uses AI to run GKnect. Coming next: AI without guardrails is a compliance lawsuit waiting to happen.

Read More Articles