(transparency logs might be a topic related to this.)
(Atproto: might be a vessel of this. And was the card on the board that got stickers for interest.)
Other things that do this? Well, formats, there are plenty:
link attestations (in-toto style) from rebuilderd
oss-rebuild does something similar to that
build info files are something like this
this list could get huge
Grab bag of thoughts (from “what are people interested in” phase):
Formats don’t explain how things are published. We’re interested in publish.
Atproto is a publishing platform.
Enumeration/discoverability is slightly distinct from distribution.
You can publish/distribute something but if it’s in a forest and nobody hears it, it can’t be enumerated/discovered.
Interest in publishing, in observability.
Some interest in trying to make more standardization of build metadata descriptions and formats.
not impressed with buildinfo files.
something more generic perhaps?
Want to talk about systems that solve lookup of name to hash. Specifically, as isolated component.
Transparency log gives more than distribution: means things can’t be removed again. Is that important?
Practically as a tool builder: a person vendors a lot of packages. Hundreds of thousands.
And one wants to get metadata about all of these things. Also later!
Want control over the firehouse of metadata.
Want a way to publish bugs and security issues to some firehose that others could consume… which is not necessarily the upstream.
(People exclaim “yes”)
Verification?
Parts
Identity
Storage
Transport Helper
(gossip?)
connect storage backends
Apps
take in some part of a data stream (only the parts they care about)
these tend to have to build indexes. search costs resources.
why do we want something other than just http servers publishing files that are signed?
sharing signatures and metadata saying where to get the file?
discoverability is a major reason.
atproto gives some things
easy to use your pds, not new infra
if you ignore the relay in atproto, you degrade to close to the “just use servers at home” thing.
atproto vs transparency logs?
discovery discovery discovery.
if i make a transparency log of my build results then i go to the world and say “please clap”.
transparency logs don’t necessarily do storage, themselves.
they just checkpoint hashes that prove existence and order of documents.
atproto doesn’t really have witnesses, necessarily? kinda?
(look more at this?, maybe it kinda does, in some parts, approximately)
do we care about atproto super concretely?
no, but… it has kinda all the right parts.
“do we all agree that package identity and metadata are distinct from data availability?”
all 8 agree.
but in subsequent discussion, yes, “data availability is a real challenge”.
to practically try to get more shared records about what’s reproducible between distros: what can we do? what subproblems are there?
rebuilderd is close to this.
in practice?
buildinfo files: do they help?
tl;dr: no.
debian and arch both have something called a buildinfo file. it is not the same.
even if it was structurally the same, it has package names. which are local to those ecosystems.
so it doesn’t actually standardize… very much at all.
about the package name agreement subproblem:
PURL?
distrotracker has some naming association between distros.
whatsrc has this name cross-mapping index for source packages.
“aren’t these things basically self-electing themselves as new name authorities?”
yes.
key point in any system we’re discussing aiming for here:
records are immutable.
let’s see this as a database schema
we can all definitely agree “output content hash” is a column that is also absolutely useful to index on.
everyone agrees on “source hash” being a column and useful.
there’s a lot of opinions about “toolchain hash” – some want to include this; defining it gets contentious.
there’s the remaining “environment” we can’t identify nor hash – kernel versions; broken chips; who knows – these tend to be found later! – these are oof.
key takeaway: seeing this as a database with indexes that can be useful is a good way to become a little more flexible how we think about this.
(does not have to literally be a database; this is just a way to think about the data’s utility and relationships)