Most AI knowledge work still gets trapped in whatever system produced it: a data catalog, a vector database, a BI tool, an agent memory store, or a pile of exported JSON.
Open Knowledge Format takes a different path.
It says: make the knowledge a directory of Markdown files with YAML frontmatter, then let people, agents, search indexes, static sites, and graph tools read the same bundle.
Open Knowledge Format (OKF) is a vendor-neutral format for representing knowledge as plain Markdown files with YAML frontmatter.
It is designed for people, agents, export pipelines, and tools to produce and consume the same corpus.
Open Knowledge Format GitHub Source Code Open Knowledge Format v0.2 Specification Original knowledge-catalog OKF Snapshot License: Apache-2.0
What is Open Knowledge Format?
OKF is a file format and convention for portable knowledge bundles.
The core object is simple:
bundle/
index.md
log.md
tables/events.md
metrics/revenue.md
references/source-doc.md
Each concept is a Markdown file. The top of the file carries YAML frontmatter. The body carries the human-readable explanation, schema, examples, links, and supporting context.
The smallest valid concept only needs a type:
---
type: Reference
---
# Some piece of knowledge
That minimalism is the point. OKF does not require a running server, schema registry, graph database, vector store, or proprietary metadata API. A bundle can live in git, a static site, an object store, a tarball, or a normal filesystem.
Repository Snapshot
I started from the URL under GoogleCloudPlatform/knowledge-catalog/tree/main/okf, but that page now says the directory is a frozen snapshot and points readers to the standalone GoogleCloudPlatform/open-knowledge-format repository.
So this analysis uses the canonical repo:
tmp/open-knowledge-format
Source state:
| Item | Value |
|---|---|
| Repository | GoogleCloudPlatform/open-knowledge-format |
| Commit | ad30107c31c06aec8a7d5636e0d1058118604e6f |
| Commit date | 2026-08-21 13:08:36 -0700 |
| Commit message | Merge pull request #6 from GoogleCloudPlatform/okf-iso-datetimes |
| Spec version | OKF v0.2 |
| Package | reference-agent 0.1.0 |
| Runtime | Python >=3.11 |
| License | Apache-2.0 |
Why OKF is Interesting
The nice thing about OKF is that it treats knowledge curation like software engineering.
A bundle is diffable. It can go through pull requests. You can review one table definition, one metric, or one policy.
You can run tests against it. You can publish it with a static site. An agent can load one directory index first, then open only the concepts it needs.
That matters for AI systems because “knowledge” is not just retrieval text. It also needs provenance, freshness, lifecycle state, and a trust signal.
OKF v0.2 makes those fields explicit:
| Signal | Where it lives | Why it matters |
|---|---|---|
sources |
Frontmatter | Records where a concept came from |
generated |
Frontmatter | Records who or what wrote the current content |
verified |
Frontmatter | Records machine or human confirmation |
status |
Frontmatter | Marks draft, stable, or deprecated concepts |
stale_after |
Frontmatter | Lets consumers warn when knowledge is out of date |
| Markdown links | Body | Build graph relationships without a graph DB |
The trust model is deliberately derived, not stored as a magic score. No verified field means unverified. Verification only by processes means machine-confirmed. Any human:<id> verifier means human-reviewed.
Tech Overview of OKF
The repository has three parts:
SPEC.md
-> OKF v0.2 format definition
src/reference_agent/
-> Python proof-of-concept producer and viewer tools
bundles/ and samples/
-> example bundles and reproducible recipes
The Python package is called reference-agent.
It uses:
| Component | Role |
|---|---|
google-adk |
Agent framework for the reference agent |
google-genai |
Gemini model access |
google-cloud-bigquery |
BigQuery dataset/table metadata source |
pyyaml |
YAML frontmatter parsing |
markdownify |
Convert fetched HTML pages into Markdown |
pytest |
Test suite |
| Cytoscape.js | Graph visualization in generated HTML |
| marked | Markdown rendering in generated HTML |
The implementation currently has one data-source connector: BigQuery. It lists datasets and tables, collapses sharded table families such as events_20240101, reads schema and partitioning metadata, and can sample rows.
Then the agent writes OKF docs. An optional web pass fetches seeded documentation pages and enriches existing concepts or creates new references/<slug> concepts.
Trying OKF Locally
This one was practical to test without containers.
Environment:
Python: 3.12.3
Root disk: 56G free, 94% used
RAM: about 50GiB available
Swap: 8.0GiB
Commands:
git clone --depth 1 https://github.com/GoogleCloudPlatform/open-knowledge-format.git tmp/open-knowledge-format
cd tmp/open-knowledge-format
python3 -m venv .venv
.venv/bin/pip install --index-url https://pypi.org/simple/ -e .[dev]
.venv/bin/pytest
Result:
39 passed in 0.34s
I also generated a static graph viewer from the included GA4 sample bundle:
.venv/bin/reference-agent visualize \
--bundle bundles/ga4 \
--out /tmp/okf-ga4-viz.html \
--name "GA4 OKF smoke"
Result:
Wrote 9 concept(s), 8 edge(s), 47911 bytes -> /tmp/okf-ga4-viz.html
That is a useful field test because it exercises the consumer side without cloud credentials.
I did not run the BigQuery/Gemini enrichment flow because that requires Google Cloud credentials, Gemini credentials, and can bill the configured Google Cloud project.
No Docker commands were used.
Running the Reference Agent
The README uses this install path:
python3.13 -m venv .venv
.venv/bin/pip install --index-url https://pypi.org/simple/ -e .[dev]
The project metadata says Python >=3.11, and the local test worked with Python 3.12.3.
To produce a bundle from BigQuery:
.venv/bin/reference-agent enrich \
--source bq \
--dataset <project>.<dataset> \
--web-seed-file <path/to/seeds.txt> \
--out ./bundles/<name>
Useful flags:
| Flag | Use |
|---|---|
--no-web |
Skip the web documentation enrichment pass |
--concept tables/events_ |
Iterate on one concept |
--web-max-pages N |
Cap web pages fetched |
--web-max-depth N |
Cap crawl hops from seed URLs |
--web-allowed-host |
Allow an additional host |
--web-allowed-path-prefix |
Restrict crawl paths |
--web-denied-path-substring |
Block noisy paths such as login pages |
Visualizing a Bundle
The viewer is the part most people can try first.
.venv/bin/reference-agent visualize --bundle ./bundles/ga4
That writes viz.html inside the bundle by default. The generated page gives you:
- a graph of concepts and Markdown links
- a detail panel for frontmatter and body content
- backlinks
- search
- type filtering
- switchable graph layouts
- badges for status, trust tier, and staleness
It is not a server. It is one static HTML artifact you can open in a browser or host with any static file server.
Where OKF Fits
OKF is not a replacement for a database catalog, schema registry, BI semantic layer, or vector database.
It is a portable exchange layer for the knowledge around those systems.
Use OKF when you want:
- a git-native knowledge corpus
- agent-readable metadata without a custom SDK
- Markdown docs with queryable frontmatter
- provenance and trust signals near the content
- static graph browsing
- a format that can survive tool changes
It is probably not the right first tool if you need a full catalog UI, automated lineage engine, permissions system, search backend, or hosted metadata platform.
OKF can feed those systems, but the spec itself is intentionally smaller.
Attested Computations
The most interesting part of v0.2 is Attested Computation.
This lets a concept say: here is the sanctioned computation, here are the allowed parameters, here is how to execute it, and here is how to verify the receipt.
That distinction is useful for agents. You do not want an LLM to improvise revenue SQL every time someone asks a question.
You want it to call the approved computation, bind declared parameters, and show a result only if the deterministic attester accepts the receipt.
In OKF terms:
verifiedchecks that the definition still matches policy- attestation checks that a specific run used the sanctioned computation
That is a practical pattern for AI analytics, data catalogs, finance metrics, and internal operations where “the answer” is not enough. You need to know how it was produced.
Conclusion
Open Knowledge Format is worth watching because it is boring in the right way.
It uses Markdown, YAML frontmatter, directories, links, timestamps, and git-friendly files. Then it adds just enough structure for agents to reason about provenance, trust, freshness, lifecycle, and computation.
The reference agent is still a proof of concept, and the current producer path is BigQuery-heavy. But the format itself is broader than Google Cloud.
You can write OKF by hand, generate it from another catalog, publish it as static files, or load it into an agent context window one concept at a time.
For self-hosters and data teams, the interesting question is not “can I deploy OKF?”
It is “which knowledge do I want to make portable before it gets trapped in another tool?”
FAQ
Is OKF a self-hosted app?
viz.html viewer, but there is no first-party Docker Compose application in this repo.
What happened to the knowledge-catalog OKF folder?
GoogleCloudPlatform/knowledge-catalog/tree/main/okf directory now identifies itself as a frozen snapshot. The canonical home is GoogleCloudPlatform/open-knowledge-format, which is what this post analyzes.
Does OKF require Google Cloud?
Can I use OKF with Obsidian or MkDocs?
What is the easiest OKF command to try first?
Run the static viewer against an included bundle:
.venv/bin/reference-agent visualize --bundle bundles/ga4
That does not require BigQuery or Gemini credentials.
Comments