Dark data vs. dark context

Dark data vs. dark context

Two kinds of invisible, one word apart

Dark data is information you collected and never used. Dark context is knowledge you never collected at all: it lives in people, not systems. They rhyme, and they are not the same thing.

Logs you never openedKnowledge no log holdsData at rest, in a warehouseJudgment at rest, in a headCaptured, then ignoredNever captured at allA discovery problemA capture problemTurn the lights on what you haveGet it out of heads before it leaves Logs you never openedKnowledge no log holdsData at rest, in a warehouseJudgment at rest, in a headCaptured, then ignoredNever captured at allA discovery problemA capture problemTurn the lights on what you haveGet it out of heads before it leaves
00 / The short answer

Same blind spot, two different causes

Dark data is data your organization has, collected and stored and then left unused. Dark context is knowledge your organization has never captured at all: it was never written down, so it sits in no system, only in people.

The quickest way to tell them apart: you can connect every database you own to an AI and still not give it your dark context, because your dark context was never data to begin with.

01 / Dark data

Collected, then ignored

The term dark data was coined by Gartner: the information assets organizations collect, process, and store during regular business activities, but generally fail to use. Server logs, old sensor readings, form submissions, call recordings, scanned documents. Captured by default, analyzed by almost no one.

Gartner estimates 80–90% of enterprise data
is dark: stored, and never used.

It is a data-management problem. The information already exists inside your systems; you are simply not looking at it. The fix is discovery and analysis, turning the lights on data you already have.

02 / Dark context

Never captured at all

Dark context is what an organization knows but never wrote down — the situated context a general AI model cannot see. The workaround everyone uses, the reason a rule exists, the client you chased and why it failed. It sits in no document because it was never recorded: it lives in people, in habits, and in the space between meetings.

It is not a data problem, it is a capture problem. There is nothing to "discover" in a system, because it was never in a system. The fix is getting it out of heads and into a shared, current form, before the person who holds it leaves.

03 / Side by side

Where each one lives

Fig. / In your systems, vs. in your people

Dark data

captured, then never used

Lives in your systems. It was recorded, it just goes unread. Anyone with access could find it; no one looks. An AI can use it the moment you connect and index it. The fix is a discovery job.

Dark context

never captured at all

Lives in your people. It was never recorded, so there is nothing to connect. Only the person who holds it knows it, and an AI stays blind to it until someone puts it into words. The fix is a capture job, and it walks out the door when they leave.

04 / The common mistake

Why "just connect your data" doesn't fix it

The usual move is to point an AI at everything you have: retrieval over the wiki, the drive, the tickets, the warehouse. That surfaces your dark data, and it is worth doing. It does nothing for your dark context, because the context was never written down to be retrieved.

You can index everything you wrote.
You cannot index what you never wrote.

A perfect search index over every document you have ever recorded still cannot find what no one ever recorded. That is the gap a general model cannot close on its own, and the gap connecting more data will never reach.

05 / The newer mistake

Captured completely, still blind

The mistake has a newer version. It no longer says connect the data you already have. It says record more of it, at higher fidelity, from further upstream. Modern observability captures every session with no sampling: device state, the order someone tapped things, network conditions, the sentiment in the review they left afterwards. The vendors selling it now argue that your bottleneck is context rather than the model.

They are right about the diagnosis and wrong about the reach. Instrumentation captures what a system emits. A decision emits nothing.

You can record every session.
You cannot record what was never emitted.

The crash log holds the failure. It does not hold the meeting where someone said this flow will confuse people and was overruled, or the reason the retry limit is three, or which customer you decided to disappoint. Push the fidelity as high as it goes and the arithmetic does not move: all of the emitted signal is still none of the unemitted.

06 / Which is your problem?

Most organizations have both

If the knowledge already exists somewhere in your systems and no one uses it, that is dark data, a discovery job. If the knowledge exists only in someone's head and no system holds it, that is dark context, a capture job.

Only one of them walks out the door when someone leaves.

07 / Common questions

Dark data and dark context, answered

Is dark context just dark data?

No. Dark data was captured and stored, then left unused: it exists in your systems. Dark context was never captured at all: it exists only in people. The same blind spot, arrived at two different ways.

Is dark data the same as dark context?

They are siblings, not synonyms. Dark data is a data-management problem, information you have but do not use. Dark context is a capture problem, knowledge you never recorded in the first place.

Can an AI use dark data?

Yes, once you surface and connect it: that is exactly what data discovery and retrieval do. An AI cannot use dark context, because dark context was never written down to be connected.

Does full observability capture dark context?

No. Observability captures what your systems emit: sessions, crashes, click order, network conditions. Dark context was never emitted by a system, because it never entered one. Recording every session still records none of the reasoning behind them.

How do you capture dark context?

Through structured conversation that gets a group's tacit knowledge into shared, current language, and then keeps it fresh as the work changes. Not a documentation sprint, and not a tool that decides for you.

Who coined the term "dark context"?

Ola Möller and Andriy Zhukov, at Dark Context, by analogy to dark matter, the unseen mass that holds things together, and as a deliberate nod to Gartner's "dark data".

darkcontext: the part you never wrote down

darkcontext:~$which one is your problem?

A short call, not a sales pitch. Bring one process that runs on what someone never wrote down — the part no dataset holds — and we'll work out whether that dark context is worth surfacing, and what surfacing it would take.

book a call

darkcontext:~$not ready to talk?

Leave an email instead. No pitch, just the occasional note as we map where dark context shows up and what helps bring it to the surface.

>