9.7 KiB
The snapshot model
A snapshot is the file zfin snapshot
writes to history/<date>-portfolio.srf, and the thing every
history- and comparison-oriented command reads back. This page explains
what a snapshot is for, which is also an explanation of what it
deliberately leaves out. For the day-to-day workflow, see
Snapshots and history.
A snapshot is a valuation record
A snapshot answers exactly one question:
What was this portfolio worth on date D, decomposed by symbol, account, and tax type?
That is the whole contract. A snapshot is not a faithful record of your lots. It records enough per-lot detail to decompose a valuation, and stops there.
This is a deliberate choice rather than an oversight, and the rest of this page is the reasoning -- partly so the decision is auditable, and partly so the same field-by-field extension proposals don't have to be re-litigated every time someone notices something missing.
Two layers, and which one wins
A snapshot file has two layers that are computed independently, and it matters which you read:
- The totals layer -- the
total,tax_type, andaccountrows. This is the authoritative valuation. It comes fromvaluation.portfolioSummary, the same code path that produces the headline numbers inzfin portfolio, and it includes portfolio-level adjustments. - The lot layer -- the
lotrows. This is a decomposition aid. It comes from a separate, direct walk over the portfolio's lots, and it deliberately carries no portfolio-level adjustments.
When the two disagree, the totals layer wins. See "Known shortcoming" below for the one case where they measurably do.
Why not a faithful lot record?
Because zfin already has one, and it is better at the job.
The git history of your portfolio.srf is the faithful lot record.
It has complete fidelity (every field of every lot, not a projection of
some of them), and it has true temporal resolution (one revision per
edit, rather than one file per day).
zfin contributions is built
directly on it: it walks git revisions, deserializes complete lots, and
matches them across revisions by a composite key.
Making snapshots into a second faithful lot record would create two records of lot identity with different cadences, different completeness, and different provenance -- and therefore an open-ended obligation to reconcile them whenever they disagree. Nothing currently asks for that.
There is also a harder lesson embedded in the contributions code. Even with the complete lot available, matching lots across time is unreliable enough that it needs a fuzzy fallback key alongside the strict one, because ordinary bookkeeping breaks strict identity: certificate-of-deposit auto-renewals rewrite the open date, account renames change the account, and reconciliation tweaks move the open price. Cross-time lot matching is a genuinely hard problem that the git-revision pipeline solves with real machinery. A daily valuation file is not going to solve it as a side effect of carrying one more column.
Proposals that were considered and rejected
If you are here because a field you want is missing, check this list first.
open_date
The strongest-looking candidate, because without it a snapshot's lots have no temporal identity at all, so they cannot be matched to lots in any other snapshot.
Rejected, because it doesn't actually deliver that. open_date is one
component of the strict lot key; a snapshot carrying it would still
lack everything the fuzzy fallback needs when the strict key breaks --
which, per above, is routine. The result would look like it enabled
cross-snapshot lot matching while quietly not doing so, which is worse
than the honest absence. If cross-snapshot lot keying is genuinely
needed, the answer is to extend the git-revision pipeline.
There is one non-identity argument for open_date: holding-period
analysis, such as decomposing unrealized gains into short- and
long-term at a past date. cost_basis is already emitted, so
open_date is the only missing input. This is a legitimately
valuation-shaped use, and it is the argument to make if the field is
ever wanted -- but it is currently hypothetical. zfin has no
holding-period logic anywhere, and tax-loss harvesting is deliberately
hand-declared in accounts.srf rather than computed from lots. Adding
a field for a consumer that does not exist is how formats rot.
split_factor
A stock lot's shares is written raw (as transacted) while its value
is split-adjusted, so the two are not related by the obvious
shares * price identity. It is tempting to emit the split factor to
close that gap.
Rejected, because it carries no information. All three quantities are already recoverable by algebra from what is emitted:
raw shares == cost_basis / open_price
effective shares == value / price
split_factor == value / (shares * price)
compare.aggregateSnapshotStocks uses the second of these and
documents it. Emitting the split factor as well would introduce a
second route to effective shares within a single record: identical
whenever the writer is correct, divergent exactly when it is buggy, and
with no rule for which one a reader should trust. That is the same
two-sources-of-truth problem as the lot-record proposal, in miniature.
The "it would let us validate the invariant" argument does not survive either: a check comparing the writer's output against the writer's own inputs, assigned in the same function moments earlier, cannot detect a wrong split factor, a wrong price, or a wrong share count. It can only catch a typo in a single expression, which is a unit test's job.
drip
Rejected as a category error. zfin contributions classifies dividend
reinvestment by reading the drip flag from git revisions of
portfolio.srf; it never reads snapshots at all. Adding drip here
would not affect that classification, or anything else.
Known shortcoming: the covered-call gap
sum(stock lot.value) can slightly exceed total::liquid in the
same file.
The cause is the two-layer split. When you hold an open, in-the-money sold call, the totals layer caps the covered underlying's market value at the option's strike price -- the shares are effectively committed at that price, so valuing them at the higher market price would overstate the portfolio. The lot layer applies no such cap; each lot is marked at plain market value.
Consequences, in ascending order of obscurity:
total::liquidis correct and remains the figure to read.- Summing lot values yourself will overstate the liquid total whenever
such a call was open on the snapshot date. The overstatement is
bounded by
(market - strike) * covered_sharesacross affected underlyings. history.aggregateSnapshotAllocationsderives per-symbol weights by dividing summed lot values by the totals-layer liquid figure, so those weights can sum to slightly more than 1 in the same circumstance.
This is documented rather than fixed because it is immaterial in
practice: it requires an open ITM sold call on the snapshot date, and
the affected consumer is an allocation weighting whose downstream use
tolerates the error. Note also that it is not diagnosable from a
snapshot alone -- deciding whether a given gap is a legitimate
covered-call cap or a writer bug requires the option's strike, which
the format does not emit. Emitting strike and multiplier is
therefore the one field addition with a concrete, valuation-shaped
consumer, should this ever need fixing properly.
Wire-format compatibility
The format is SRF, and its compatibility behavior is what makes the format safe to extend:
- Default-valued fields are elided on write. A field equal to its default does not appear in the file at all.
- Unknown fields are ignored on read. A reader skips fields it doesn't know about.
- Together these give compatibility in both directions: old readers tolerate new files, and new readers tolerate old files.
Two constraints follow for anyone extending a record type:
- Every field must have a default. Fields are matched by name and absent ones are filled from their default; a field with no default makes every previously-written snapshot fail to parse.
kindmust stay first. It is the union discriminator, and the reader requires it as the first field on the line.
snapshot_version and ambiguous absence
Each snapshot carries a snapshot_version, currently 1. Nothing
reads it to gate behavior, so bumping it is normally documentary.
There is one case where it becomes load-bearing. Because default-valued
fields are elided, a new field whose value happens to equal its default
is absent from the file -- and a reader cannot tell that apart from "the
writer predates this field entirely." Bump snapshot_version when,
and only when, a new field's absence is semantically ambiguous.
Novelty alone is not a reason.
Worked examples: a drip: bool = false or an open_date: ?Date = null
would each require a bump, because absence could mean either "false /
unknown" or "old writer." A split_factor: f64 = 1.0 would not, since
1.0 means "no adjustment" either way.
Writers and readers are separate deployments
Snapshots are often written by a scheduled job and read by interactive
commands, and those two can be running different builds of zfin for
weeks at a time. So a newly added field appears only in files written
after the writer's build is refreshed, and never appears in files
already on disk. zfin snapshot never reads existing snapshots, so an
older writer is safe; just don't assume a field's presence based on the
version of the binary you happen to be reading with.