Gothenburg 2026 - Day 2 - Transparency logging
transparency-logging
why to use transparency logs for rb:
- make lies something that people can detect and monitor
- the results of rebuilderd, what distribution, search, merge
log backed partition maps patricia merkle tree
if you transparency log is a set of keys that can be partitionned,
for example by the partition key, and you can that into an index of this
the key can be triple (the package name, version number, and the architecture) for one triple, there are many binaries,
what are the log entries? -> a hash of a binary? not only, we need the name, something else?
content of a potential entry: - who is the rebuilder, - svc / source / package name, - version,
then:
arch1,binary1 -> hash1
arch2,binary2 -> hash2
we need the hash of the source in here or not? - the hash of the DSC file?
we wish this could be compatible between distros, however they have different schemes for source, package name, version schemes… etc – and if different distros use the same log, how do we tell?
there are different conventions on how to name things :o) and different conventions on how to hash sources
why not having a long string, that different people fill, and grow from there?
[scrhash (as per whatsrc) ] ->
<rebuilderID, target(Architecture, etc), timestamp-of-the-build,
<resultLabel -> resultHash>>
Maybe instead of a “things that looks like a map” it should be maybe “things that look like rows in a database”. This would let us choose what properties we want fast lookups for later.
How to do it, if you want to use sigsum?
The “rich entry” that contains all the infos (the “row” with metadata) will be hashed / a checksum, and the monitor tail the log, get these checksums, and go to an object storage / database that does checksum -> rich entry
Side comment: good practice in adding transparency logs, is to avoid people writing “simpler clients” that discards burden of thinking about transparency logs. One way to do this, is to do the “content resolution” from leaf index (entry id / number of the leaf) to the content.
What about poisoning?
The fields of the entries (like package name) can be poisoned by the rebuilders that are on the allow-list of submitters of the logs. One rebuilder that gets broken, would be able to poison the log of everyone else.
Where we say “poisoned” here we mean: content that would later want to be removed. (Don’t make us provide examples. PII would be the most polite one.)
The minimum viable thing
a record format that contains the build information, that contains the things that matters for rb
that records get pushed in a content adressed storage (and package maintainers have the option to remove poisoned records from the storage if necessary, since it’s not directly in the log)
finally, you have in the log only the hash (of the hash) of the record format, and that protects the log from being too exposed to poisoned submissions by evil rebuilders
and to make it useful, is there is a need for an index service, for performance. This can be like a postgres database that does search (there is some trust to that database but the database can be discarded or rebuild by anyone), or it can be something more complicated but with better trust properties like a Verifiable Log Backed Map.
one thing we didn’t talk about yet
when you submit the hash of the signed record format
you get a "proof of logging" / a proof of publication
the verifier previously, was verifying:
- the signature (a rebuilderd did the rebuild)
the verifier would like to verify additionnaly:
- the proof of publication
because I want to know that a rebuilder was forced to publish a result (so that if it’s a bad rebuilder and trying to trick me, there is public proof of this i can point at afterwards).
after more discussion: this is an option but getting a “proof of logging” at submission time and carrying it around forever is not necessary. This can also be solved by inclusion proofs in the log generated at any later time. (And if someone will be checking more than one entry, which will probably be typical for build record inquiries (since we’d be comparing them), getting the inclusion proofs later will tend to be simpler and more efficient.)
summary of the notes, on the submit path:
- build !
- submit BuildRecord to the log store full record in a DataPile – it’s a CAS (log stores H(H(BuildRecord)) only)
- mantain LookupService (for speed and usability) example: for example, it maps source-hash (or whatever, choose later) to… <—what-to-put-here—>*
The instinct for <—what-to-put-here—> is the checksum of BuildRecord
However!
That enables the LookupService to invent entries that don’t exist in the log, if LookupService colludes with DataPile. A better picture is
<---what-to-put-here--->
to be the “leaf index” (number of the entry in the log) so it forces: 1. go to LookupService, ask about srchash or packagename, get leaf-index 2. go to log with leaf-index, ask about it, get H(H(BuildRecord)) as result 3. go to DataPile and learn BuildRecord
This makes sure that the transparency log is load-bearing during the lookups, which addresses big picture issues like “what if people start to trust the LookupService too much in practice, and the LookupService thus becomes able to fool people in ways the transparency log is meant to prevent”.
On the read path:
R1. Check lookup table for whatever property is interesting. Gives you a log entry index. R2. Look up the log entry index in the log. (Get log checkpoint & check witness sigs & check inclusion proof.) Congrats, now you have H(H(BuildRecord)), verified. R3. Fetch full BuildRecord from DataPile, using H(H(BuildRecord)). (Double check that whatever property you searched for in LookupService is actually as it should be in the BuildRecord.)
With this journey, we know and verified:
- The BuildRecord we searched up was publicly logged. Everyone can agree to this.
- The LookupService has no ability to lie about facts in the BuildRecord.
There’s one thing that this doesn’t make trustable and verified:
- The LookupService can falsely say there are no matches to your search. (It can’t make up entries, but it can quietly not tell you about real ones.)
The solution to this is simply: okay, then don’t trust the LookupService: but you can rebuild the indexes that the lookup service offers, directly from the log.
Alternatively, this is where VLBM’s are more incrementally verifiable and can give better trust than just plain “throw postgres at it”.