At the end of August a customer posed the question to my team: how can we use OKF with or instead of Foundry IQ and Fabric Ontologies for our agentic AI use cases?

Customer questions like this is the thing I like most about my job as a Solution Engineer at Microsoft because there isn’t an answer you can just look up - you have to go figure it out which means you’re about to learn something new and with any luck, interesting.

In this case it gave me a reason to think about using OKF in the real world, and given that I’ve been managing my life as Markdown files tracked in Git since roughly 2014 and have been a Obsidian user since about 2021, this is an area I have certain opinions about. Having gone through the exercise of looking at how OKF and Microsoft IQ could fit together, I came away with a critique of OKF that I would like to share.

Anticipation

After reading some OKF headlines I was pretty excited. As a heavy Obsidian user with two multi-thousand-node vaults my first thought was that maybe this will give me better standards for how I structure knowledge, collections of typed nodes and relationships between nodes.

Graph view of my personal Obsidian vault
My 5+ year old personal Obsidian vault contains 11,882 Markdown notes and about 45K other files.

Wait, that’s it?

But about 20 minutes into reading the standard and crafting my first OKF bundle to support a demo for my customer I realized - there’s not much here.

To be exact my thoughts were: so this is just an Obsidian vault (or your choice of tool or convention for organizing Markdown files) with a mandatory type front matter key but no actual type system, some recommended meta data for concerns such as provenance and an optional okf_version: "0.2" declaration on index.md in the root. That’s it?

Having built large Markdown based knowledge graphs, I think there are a number of areas for further improvement.

Schema

OKF defines the mandatory type key but does not provide any mechanism to define what that type means or the data that should be included in a node of that type.

This is one of the main things I was hoping to find as I started to read the spec because i already have my own way of doing typed notes in Obsidian. For example, I have a collection of country notes that I can use whenever I need to refer to a country from another note - here’s the Thailand.md country note from my vault:

---
api: "bren/geo/v1"
kind: "country"
aliases:
  - TH
  - TH 🇹🇭
iso_2: TH
iso_3: THA
flag: 🇹🇭
name: Thailand
in_regions:
  - "[[South-East Asia]]"
  - "[[Asia]]"
world_bank_page: "https://data.worldbank.org/country/thailand"
gdp_value: 514
---

The Kubernetes-ish api and kind keys are what I chose to represent the namespaced type declaration for a note.

My vault is replete with these sorts of typed notes, but I don’t have a schema system to lean on. I was hoping to get one from OKF but they explicitly put that out of scope.

Namespaces and Composability

In the olden times we needed an open standard format for structured data and so we invented XML, named after our generation, Gen X. Later Gen Y came along and couldn’t handle all the keystrokes so they created YAML. It’s fine I guess, but it melts down if a space is missing which XML stoically would never do.

Anyway, to define the structural and value constraints for a vocabulary in XML, we created XML Schema. An XML document can combine vocabularies from multiple namespaces, with prefixes as aliases for namespace URIs and xsi:schemaLocation hints for the corresponding schema documents.

<catalog:countryRecord
    xmlns:catalog="https://example.org/vocab/catalog"
    xmlns:geo="https://example.org/vocab/geography"
    xmlns:meta="https://example.org/vocab/metadata"
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
    xsi:schemaLocation="
      https://example.org/vocab/catalog country.xsd
      https://example.org/vocab/geography geography.xsd
      https://example.org/vocab/metadata metadata.xsd">
  <meta:title>Thailand</meta:title>
  <geo:iso3166Alpha2>TH</geo:iso3166Alpha2>
  <catalog:category>country</catalog:category>
</catalog:countryRecord>

This mechanism had a some important consequences:

  • You can disambiguate named types from different vocabularies by their namespace - so for example the country types from two different vendors can be distinguished or even co-exist cleanly
  • You can compose different vocabularies in a single XML document.
  • A piece of software that understands one schema but not others that may have been composed into a document is able to pull out just the bits it knows without worrying about the rest.

I was hoping to find a nice namespacing / composition mechanism in OKF but none is offered.

Semantic Graph

This is the biggest gap given that we are talking about knowledge graphs, and given the prior art that already offers solutions for this.

OKF supports links in the prose (the text / HTML body). As the spec says, this creates untyped links but you can understand their meaning by reading and interpreting the prose around them:

A link from concept A to concept B asserts a relationship. The specific kind (parent/child, references, joins-with, depends-on) is conveyed by the surrounding prose, not by the link itself. Consumers that build a graph view typically treat all links as directed edges of an untyped relationship.

This is fine but ideally we also want a way to define links that have an intrinsic semantic meaning. For example my country sample above has the following:

in_regions:
  - "[[South-East Asia]]"
  - "[[Asia]]"

This is using the linked properties feature that was added to Obsidian a couple of years ago to create relationships between my Thailand node and two other nodes but importantly the relationships have semantic meaning: in_region. This is effectively a RDF triple:

Subject → Predicate → Object

Thailand → InRegion → SouthEastAsia

This is meaningful and traversable without the cost and uncertainty of LLM inference to interpret the relationship.

In OKF the Thailand country node might look like this:

---
type: "country"
---
Thailand is a country in [South-East Asia](South-East Asia.md) in greater [Asia](Asia.md)

While this can be read and understood by a LLM navigating the wiki by progressive disclosure, it can’t be traversed efficiently and precisely.

OKF leverages the idea of progressive disclosure. That means your LLM gets a summary or overview of the content you’re offering to it, and based on that it can selectively request the full details of the content elements that it thinks are relevant to it’s goal. This is a strategy for saving tokens and context window space by loading only what is relevant. It’s the strategy used to selectively load agent skills.

In the case of agent skills, it makes sense: we start with a list of the descriptions of all available agent skills, and once the LLM has decided which skill should be utilized the full definition of the skill is loaded. It’s one level of progressive disclosure.

Applying this to navigating a knowledge graph however is questionable, because any non-trivial lookup is going to be multi-hop, burning tokens and milliseconds at each hop.

In most if not all cases we will want to query (eg with SPARQL) or traverse (eg with Gremlin) the graph to resolve the nodes we want the LLM to then get as context.

OKF does define some “path-value fields”, and defines a “references convention” which uses front matter elements to define relationships. These are specific relationship types built into the spec that could have been defined as typed relationships on a general RDF-based framework.

Of course you’re not prevented from creating RDF-style relationships in the YAML front matter but it’s left to you to figure out how that should work.

Standardized Meta Data

OKF proposes standard YAML frontmatter elements for describing:

  • Provenance
  • Trust
  • Lifecycle and Freshness
  • Attested Computation

It’s debatable whether these elements are truly universal enough that they should be baked into the standard. It’s also debatable whether these exact definitions of provenance, lifecycle, etc are suitable for all use cases which they need to be as part of the standard.

But even if they are that universal - if we had a good typing system with namespacing and composability then they could be defined as composable types that could be selectively added where needed without needing to be a concern of the framework. In fact you could tailor your own provenance model to meet your local needs or augment one provided by the OKF project.

Conclusion

There may be good reasons why OKF is defined in the minimal way that it is, and maybe my critique is arising from my own hopes and dreams that OKF was never meant to fulfill.

At the moment though - in my view it misses the mark for what we need for a Markdown-based knowledge graph standard given that it has no type system, composition mechanism or semantic graph. Also because progressive disclosure will be slow and expensive on real-world multi-hop traversal, and because it provides less than what people in the second brain / Obsidian / Notion world are already doing.