All work

Sextant

A mini-Foundry for maritime data. It maps messy records onto an ontology, merges duplicates into real entities, and can prove where every single value came from.

When
Jun–Aug 2026
What
Self-led project
Stack
C++20, React, TypeScript, CMake
Links
Code on GitHub
# cluster, fuse, write entities with provenance
$ sextant resolve
  1451 source records, 470 candidate pairs
  1187 entities written, 4201 properties,
  4201 provenance records
  dedup ratio 0.1819  (1451 records -> 1187 entities)

# the lineage round trip
$ sextant explain
lineage round trip
  1187 entities, 4201 properties
  4201 verified, 0 failed  ->  100.00%

# a quarter of arrivals at Rotterdam
$ sextant query --type Port --where locode=NLRTM \
    --link arrivals --from 2026-04-01 --to 2026-07-01
  24 result(s)
  cost
    index_used           TIDX
    keys_scanned         25
    bloom_rejections     0
    elapsed_us           347

Why it exists

A data platform is only trustworthy if you can point at any number and ask where it came from, and get an answer you can check. Most systems keep that history as a log or a comment. In Sextant it is an invariant that the system tests.

The data is real and messy: port and vessel records from the NGA World Port Index (CSV), UN/LOCODE (CSV) and Finland’s Digitraffic service (REST/JSON, including AIS ship positions). The same port shows up under different names, codes and coordinates in each source.

The name comes from navigation. A sextant fixes your position by combining several independent observations, which is what entity resolution does.

What I built

A storage engine underneath

An LSM-tree key-value store with a write-ahead log, memtable, SSTables, bloom filters, a block cache, leveled compaction and crash recovery. I chose to build it instead of linking RocksDB, and I followed LevelDB’s design closely, so this part is a reimplementation for learning rather than a new design.

Everything on top is mine

  • Ontology. A declarative schema over twelve keyspaces, with connectors for all three sources.
  • Entity resolution. Blocking, pairwise scoring, veto-constrained clustering and fusion. 1,451 records become 1,187 entities.
  • Cell-level lineage. Every property stores the raw source row and the chain of transforms that produced it.
  • Query planner. It returns its plan and its cost (keys scanned, blocks read, bloom rejections, microseconds) with every answer.
  • Interface. An HTTP API and a React front end with a lineage drawer, a link graph and a review queue.

Results

of 4,201 properties across 1,187 entities replay from their source row
100%
F1 resolving 1,451 records into 1,187 entities
0.991
batched writes per second at 1.20× write amplification
1.53M
tests passing on Linux, Windows and macOS, clean under ASan, UBSan and TSan
418

sextant explain walks every property of every entity, fetches the raw source row it names, re-applies the transform chain and checks that the result equals the stored value. It has a negative control, because a check that cannot fail is not a check.

Vetoes in the clustering step lift precision from 0.973 to 1.000. A three-month window query is answered in 359 µs over 25 keys, where a full scan walks 92.

What the numbers don’t say

The storage figures are microbenchmarks on synthetic keys on one machine. 1.53M writes per second is one process writing batched 100-byte values into a warm page cache. With sync=true the same engine does about 655 writes per second, which is the fsync floor, so the batched number is not a durability figure.