Binaries

Facts can describe a payload, but JSON is not an efficient container for its bytes. Base64 increases payload size and puts large values inside JSON processing. Large media can also dominate queries, patches, and history when its bytes travel with ordinary facts. Binary storage keeps those bytes out of ordinary facts. An object entity connects a payload to durable knowledge.

Content identifies bytes

A SHA-256 digest identifies one exact byte sequence independently of its file name, location, entity, or domain meaning. Equal payloads therefore use one content-addressed stored payload. Different facts can give those shared bytes different roles without copying the payload. A changed byte sequence receives a different digest and remains a different immutable version.

In DUST, a digest value uses the #sha256 sigil:

dust82B
digest {#sha256 e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855}

Only the digest supplies payload identity. File names and directories are object facts, not part of that identity. Applications use other facts for ownership, titles, lifecycle state, relationships, and other domain meaning.

Chunks follow content

Stardust keeps each payload as an ordered list of chunks. A manifest identifies the chunks by digest and records their order. Each chunk has its own SHA-256 digest. A rolling hash over the most recent 64 bytes sets chunk boundaries. Chunks have a 256 KiB minimum, a 1 MiB average, and a 4 MiB maximum size. An insertion or deletion therefore moves nearby boundaries without shifting every later chunk.

Two older structures share this idea. A Merkle tree identifies data blocks by hash and combines their hashes into one root. Equal content has equal identity, and a change affects only its blocks.

A Prolly tree is the probabilistic B-tree behind Noms and Dolt. It adds content-defined boundaries. A rolling hash decides where each node ends instead of a fixed position. An edit changes only nearby nodes, and equal data always forms the same tree.

A Stardust payload is close to one level of that design. Content sets the boundaries, so two file versions can share chunks away from a change. Each chunk has a hash, so equal chunks occupy one place in the store. The manifest is a flat list. The payload digest is the SHA-256 of the complete byte sequence.

Payload manifests and chunks live in the same embedded store as facts. Stardust does not require a separate object store. Published manifests and chunks are never deleted, because every historical fact that names a digest must stay readable.

Objects are files and entities

The mount exposes binary data under objects/. mkdir creates a directory, and closing a new file publishes its complete payload. The mount commits new file and directory metadata 20 ms after the last create. Pending files stay out of path lookup and directory listings until that commit. Pending directories remain traversable during a folder copy. A successful close or fsync can precede the commit, so a crash can lose pending creates.

Per-file fsync leaves new files in the pending batch, which lets a copier group them. Overwrites commit when their file handles close. Rename and unlink commit at once. Reads return only the requested byte range, so a large payload is never fully buffered for one read.

Each directory or file under objects/ is an entity. A file entity holds its name, its parent, and the stardust/object/digest field that names its payload. A directory holds no digest. The user.stardust.digest extended attribute exposes a file's digest as lowercase hex, and user.stardust.node exposes its entity id.

Two read-only views sit under each of format/json/data/objects/ and format/dust/data/objects/. file-to-hash/ mirrors the name tree and renders each file's digest. hash-to-files/ is a reverse index sharded by the first two digest bytes, and each leaf lists every object path that carries that digest. The leaf documents use the selected JSON or DUST format.

Publication cannot outrun bytes

A write stages its chunks until it finishes. Publication then installs the chunks and the manifest in one store transaction. A payload that already exists is reused rather than stored again. Startup removes stages that a crash left unfinished.

Every digest reference is checked against a published manifest before the object commit. A rejected commit can therefore leave an unreferenced payload, but it cannot create a reference to unpublished bytes.

Removing an object retracts its facts. The object disappears from the current tree, but its earlier facts stay in history and its payload stays stored.

A #sha256 value in any other field is an independent checksum. Only stardust/object/digest facts make a payload part of backup work.

Processing intent belongs in facts

Applications often need thumbnails, transcodes, extracted text, previews, or other derived payloads. A mutable job record outside Stardust can hide why processing started and which inputs produced its result. Facts can instead record the request, source digest, parameters, processor identity, status, and result digest. Retries and alternative processors then add history without changing the source payload or erasing earlier intent.

Processing can occur outside Stardust because specialized tools can need network, hardware, or unbounded work. Commit notifications and live queries can inform those tools about committed intent, as Eventbus explains. The external processor writes result bytes under objects/ and commits a Patch that records the outcome. Durable intent and lineage remain queryable even though payload processing occurs elsewhere.

Video is one example

A video application can keep source footage as immutable payloads and describe an edit as data. An OpenTimelineIO document can identify cuts, ordering, timing, transitions, and source references without changing the footage. Rendering creates another payload whose facts connect it to the source and edit definition. Alternative edits can share source bytes while keeping their different decisions and results.

This example does not prescribe a media workflow or add a built-in video editor. Another application can use different facts, schemas, processing tools, or document formats. The same separation applies to documents, models, archives, and other binary data. Applications own processing behavior while binary storage owns exact bytes and their integrity.

History closes over payloads

Transaction history keeps every payload that a historical object fact references. A backup packet lists the payloads that its transaction's stardust/object/digest facts assert, and the backup target keeps their manifests and chunks under content-addressed paths. Restore fetches those payloads for each transaction and publishes them in the same database transaction as its facts. A missing or mismatched chunk makes the backup history corrupt rather than producing an object without bytes.

Binary payloads extend the historical model without forcing large bytes into facts. Content addressing gives exact identity, while facts record changing meaning and relationships.