Most AI knowledge work still gets trapped in whatever system produced it: a data catalog, a vector database, a BI tool, an agent memory store, or a pile of exported JSON.

Open Knowledge Format takes a different path.

It says: make the knowledge a directory of Markdown files with YAML frontmatter, then let people, agents, search indexes, static sites, and graph tools read the same bundle.

Open Knowledge Format (OKF) is a vendor-neutral format for representing knowledge as plain Markdown files with YAML frontmatter.

It is designed for people, agents, export pipelines, and tools to produce and consume the same corpus.

Open Knowledge Format GitHub Source Code Open Knowledge Format v0.2 Specification Original knowledge-catalog OKF Snapshot License: Apache-2.0

What is Open Knowledge Format?

OKF is a file format and convention for portable knowledge bundles.

The core object is simple:

bundle/
  index.md
  log.md
  tables/events.md
  metrics/revenue.md
  references/source-doc.md

Each concept is a Markdown file. The top of the file carries YAML frontmatter. The body carries the human-readable explanation, schema, examples, links, and supporting context.

The smallest valid concept only needs a type:

---
type: Reference
---

# Some piece of knowledge

That minimalism is the point. OKF does not require a running server, schema registry, graph database, vector store, or proprietary metadata API. A bundle can live in git, a static site, an object store, a tarball, or a normal filesystem.

Repository Snapshot

I started from the URL under GoogleCloudPlatform/knowledge-catalog/tree/main/okf, but that page now says the directory is a frozen snapshot and points readers to the standalone GoogleCloudPlatform/open-knowledge-format repository.

So this analysis uses the canonical repo:

tmp/open-knowledge-format

Source state:

Item Value
Repository GoogleCloudPlatform/open-knowledge-format
Commit ad30107c31c06aec8a7d5636e0d1058118604e6f
Commit date 2026-08-21 13:08:36 -0700
Commit message Merge pull request #6 from GoogleCloudPlatform/okf-iso-datetimes
Spec version OKF v0.2
Package reference-agent 0.1.0
Runtime Python >=3.11
License Apache-2.0

Why OKF is Interesting

The nice thing about OKF is that it treats knowledge curation like software engineering.

A bundle is diffable. It can go through pull requests. You can review one table definition, one metric, or one policy.

You can run tests against it. You can publish it with a static site. An agent can load one directory index first, then open only the concepts it needs.

That matters for AI systems because “knowledge” is not just retrieval text. It also needs provenance, freshness, lifecycle state, and a trust signal.

OKF v0.2 makes those fields explicit:

Signal Where it lives Why it matters
sources Frontmatter Records where a concept came from
generated Frontmatter Records who or what wrote the current content
verified Frontmatter Records machine or human confirmation
status Frontmatter Marks draft, stable, or deprecated concepts
stale_after Frontmatter Lets consumers warn when knowledge is out of date
Markdown links Body Build graph relationships without a graph DB

The trust model is deliberately derived, not stored as a magic score. No verified field means unverified. Verification only by processes means machine-confirmed. Any human:<id> verifier means human-reviewed.

Tech Overview of OKF

The repository has three parts:

SPEC.md
  -> OKF v0.2 format definition

src/reference_agent/
  -> Python proof-of-concept producer and viewer tools

bundles/ and samples/
  -> example bundles and reproducible recipes

The Python package is called reference-agent.

It uses:

Component Role
google-adk Agent framework for the reference agent
google-genai Gemini model access
google-cloud-bigquery BigQuery dataset/table metadata source
pyyaml YAML frontmatter parsing
markdownify Convert fetched HTML pages into Markdown
pytest Test suite
Cytoscape.js Graph visualization in generated HTML
marked Markdown rendering in generated HTML

The implementation currently has one data-source connector: BigQuery. It lists datasets and tables, collapses sharded table families such as events_20240101, reads schema and partitioning metadata, and can sample rows.

Then the agent writes OKF docs. An optional web pass fetches seeded documentation pages and enriches existing concepts or creates new references/<slug> concepts.

Trying OKF Locally

This one was practical to test without containers.

Environment:

Python: 3.12.3
Root disk: 56G free, 94% used
RAM: about 50GiB available
Swap: 8.0GiB

Commands:

git clone --depth 1 https://github.com/GoogleCloudPlatform/open-knowledge-format.git tmp/open-knowledge-format
cd tmp/open-knowledge-format
python3 -m venv .venv
.venv/bin/pip install --index-url https://pypi.org/simple/ -e .[dev]
.venv/bin/pytest

Result:

39 passed in 0.34s

I also generated a static graph viewer from the included GA4 sample bundle:

.venv/bin/reference-agent visualize \
  --bundle bundles/ga4 \
  --out /tmp/okf-ga4-viz.html \
  --name "GA4 OKF smoke"

Result:

Wrote 9 concept(s), 8 edge(s), 47911 bytes -> /tmp/okf-ga4-viz.html

That is a useful field test because it exercises the consumer side without cloud credentials.

I did not run the BigQuery/Gemini enrichment flow because that requires Google Cloud credentials, Gemini credentials, and can bill the configured Google Cloud project.

No Docker commands were used.

Running the Reference Agent

The README uses this install path:

python3.13 -m venv .venv
.venv/bin/pip install --index-url https://pypi.org/simple/ -e .[dev]

The project metadata says Python >=3.11, and the local test worked with Python 3.12.3.

To produce a bundle from BigQuery:

.venv/bin/reference-agent enrich \
  --source bq \
  --dataset <project>.<dataset> \
  --web-seed-file <path/to/seeds.txt> \
  --out ./bundles/<name>

Useful flags:

Flag Use
--no-web Skip the web documentation enrichment pass
--concept tables/events_ Iterate on one concept
--web-max-pages N Cap web pages fetched
--web-max-depth N Cap crawl hops from seed URLs
--web-allowed-host Allow an additional host
--web-allowed-path-prefix Restrict crawl paths
--web-denied-path-substring Block noisy paths such as login pages

Visualizing a Bundle

The viewer is the part most people can try first.

.venv/bin/reference-agent visualize --bundle ./bundles/ga4

That writes viz.html inside the bundle by default. The generated page gives you:

  • a graph of concepts and Markdown links
  • a detail panel for frontmatter and body content
  • backlinks
  • search
  • type filtering
  • switchable graph layouts
  • badges for status, trust tier, and staleness

It is not a server. It is one static HTML artifact you can open in a browser or host with any static file server.

Where OKF Fits

OKF is not a replacement for a database catalog, schema registry, BI semantic layer, or vector database.

It is a portable exchange layer for the knowledge around those systems.

Use OKF when you want:

  • a git-native knowledge corpus
  • agent-readable metadata without a custom SDK
  • Markdown docs with queryable frontmatter
  • provenance and trust signals near the content
  • static graph browsing
  • a format that can survive tool changes

It is probably not the right first tool if you need a full catalog UI, automated lineage engine, permissions system, search backend, or hosted metadata platform.

OKF can feed those systems, but the spec itself is intentionally smaller.

Attested Computations

The most interesting part of v0.2 is Attested Computation.

This lets a concept say: here is the sanctioned computation, here are the allowed parameters, here is how to execute it, and here is how to verify the receipt.

That distinction is useful for agents. You do not want an LLM to improvise revenue SQL every time someone asks a question.

You want it to call the approved computation, bind declared parameters, and show a result only if the deterministic attester accepts the receipt.

In OKF terms:

  • verified checks that the definition still matches policy
  • attestation checks that a specific run used the sanctioned computation

That is a practical pattern for AI analytics, data catalogs, finance metrics, and internal operations where “the answer” is not enough. You need to know how it was produced.

Conclusion

Open Knowledge Format is worth watching because it is boring in the right way.

It uses Markdown, YAML frontmatter, directories, links, timestamps, and git-friendly files. Then it adds just enough structure for agents to reason about provenance, trust, freshness, lifecycle, and computation.

The reference agent is still a proof of concept, and the current producer path is BigQuery-heavy. But the format itself is broader than Google Cloud.

You can write OKF by hand, generate it from another catalog, publish it as static files, or load it into an agent context window one concept at a time.

For self-hosters and data teams, the interesting question is not “can I deploy OKF?”

It is “which knowledge do I want to make portable before it gets trapped in another tool?”

FAQ