LZ76 and FlashBack compression graphs for immune receptor repertoire analysis
Documentation · Quick Start · API Reference · Report Bug
LZGraphs is a Python library that turns T-cell and B-cell receptor CDR3 sequences into probabilistic directed graphs. It ships two graph families on a shared C core:
LZGraph: built from Lempel-Ziv 76 compression. Supports V/J gene annotation, three encoding variants, and alzgCLI.FlashBackGraph: a Markovian DAG built from FlashBack tokenization (recursive run-peeling from both ends of the sentinel-wrapped sequence). Diversity, entropy, path counting, and the p-sequence spectrum have closed-form forward-DP solutions; sequence simulation is still sampled. Driven from Python or fromlzg flashback.
Both classes share a common surface for scoring, simulation, diversity, graph algebra, posterior personalization, and binary serialization. See When to use which for a comparison.
An LZGraph built from three CDR3s. @ and $ are start/end sentinels; subpattern nodes carry position suffixes.
pip install LZGraphsRequires Python 3.9 or later. Wheels are published for Linux (x86_64 and aarch64, glibc 2.28+ or musl 1.2+), macOS (Apple Silicon), and Windows (x86_64), for CPython 3.9–3.13. Release history: CHANGELOG.md.
A container image with the lzg CLI preinstalled is also published to GHCR:
docker run --rm -v "$PWD:/data" ghcr.io/mutejester/lzgraphs build /data/repertoire.tsv -o /data/repertoire.lzgFor programmatic use, all classes accept a plain list of CDR3 strings:
LZGraph(['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', ...], variant='aap')For files, LZGraph.from_file, FlashBackGraph.from_file, and the lzg CLI (lzg build / lzg flashback build) all read through the same format-detection pipeline and accept the same formats, so any of the three build the identical graph from the identical file:
| Format | Layout | Example |
|---|---|---|
| Plain | one sequence per line | CASSLEPSGGTDTQYF |
| Seq + count | sequence\tcount (tab-separated) |
CASSLEPSGGTDTQYF\t42 |
| AIRR-compatible tabular | tab- or comma-separated, with header row | junction_aa, v_call, j_call, ... |
| FASTA | > header line, then sequence line(s) |
>seq1 / CASSLEPSGGTDTQYF |
| FASTQ | 4-line records: header, sequence, +, quality |
@seq1 / CASSLEPSGGTDTQYF / + / IIIIIIIIIIIIIII |
For AIRR-style tabular input: the sequence column is auto-detected from junction_aa / cdr3_amino_acid / cdr3_aa (variant aap), junction / cdr3_rearrangement (variant ndp), or any column named sequence/cdr3/seq. Gene calls come from v_call / j_call and must use IMGT-style notation (e.g. TRBV5-1*01); FlashBackGraph has no gene-annotation model, so a gene column present in the file is simply not read. Malformed sequence fields, and AIRR rows whose productive column is present and not truthy, are dropped and counted rather than built into the graph.
Compression is detected from file content, not the filename: gzip, bzip2, and xz are supported out of the box, and zstd is supported if the optional zstandard package is installed.
A clean, uncompressed, plain (or sequence<TAB>count) file streams straight into the C builder in constant memory, which is what makes from_file suitable for very large repertoires; everything else (compressed input, a headered/FASTA/FASTQ file, or a plain file with a byte-order mark, stray carriage returns, or a malformed line near its start) is instead read and validated in Python first, exactly as lzg build does for the same file. One residual case is intentionally not covered by that fast-path check: a malformed line far past the start of an otherwise clean, very large plain file can still reach the C builder and be ingested as-is rather than dropped, since fully validating a multi-gigabyte file up front would defeat the constant-memory streaming from_file exists to provide. For LZGraph.from_file and lzg build, passing strict_input=True / --strict-input validates the whole file up front if you need that stronger guarantee; FlashBackGraph.from_file and lzg flashback build do not currently expose an equivalent flag.
from LZGraphs import LZGraph
# Build a graph from CDR3 amino acid sequences
graph = LZGraph(
['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', 'CASSLEPQTFTDTFFF',
'CASSLGQGSTEAFF', 'CASSLGIRRT'],
variant='aap',
)
# Score a sequence
log_p = graph.pgen('CASSLEPSGGTDTQYF')
print(f"log P(gen) = {log_p:.2f}")
# Simulate new sequences
result = graph.simulate(1000, seed=42)
print(f"Generated {len(result)} sequences")
# Diversity
print(f"D(1) = {graph.effective_diversity():.1f}")
print(f"D(2) = {graph.hill_number(2):.1f}")from LZGraphs import LZGraph
sequences = ['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF',
'CASSLEPQTFTDTFFF', 'CASSLGQGSTEAFF']
graph = LZGraph(
sequences,
variant='aap',
v_genes=['TRBV16-1*01', 'TRBV1-1*01', 'TRBV5-1*01', 'TRBV7-2*03'],
j_genes=['TRBJ1-2*01', 'TRBJ1-5*01', 'TRBJ2-7*01', 'TRBJ1-2*01'],
)
# Gene-constrained simulation
result = graph.simulate(100, sample_genes=True, seed=42)
print(result.v_genes[0], result.j_genes[0])| Variant | Input | Node format | Best for |
|---|---|---|---|
'aap' |
Amino acid CDR3 | C_2, SL_6 |
Most TCR/BCR analysis |
'ndp' |
Nucleotide CDR3 | TG0_4 |
Nucleotide-level analysis |
'naive' |
Any strings | C, SL |
Motif discovery, ML features |
lzg build repertoire.tsv -o rep.lzg
lzg score rep.lzg sequences.txt
lzg diversity rep.lzg
lzg simulate rep.lzg -n 10000 --seed 42
lzg compare healthy.lzg disease.lzgfrom LZGraphs import FlashBackGraph
# Build a Markovian DAG from CDR3 sequences
graph = FlashBackGraph(
['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', 'CASSLEPQTFTDTFFF',
'CASSLGQGSTEAFF', 'CASSLGIRRT'],
)
# Score a sequence (exact forward DP, no MC)
log_p = graph.pgen('CASSLEPSGGTDTQYF')
print(f"log P(gen) = {log_p:.2f}")
# Simulate from the Markovian distribution
result = graph.simulate(1000, seed=42)
# Diversity, entropy, path count: closed-form via forward DP
print(f"D(1) = {graph.effective_diversity():.1f}")
print(f"D(2) = {graph.hill_number(2):.1f}")
print(f"# distinct paths = {graph.path_count:,}") # exact int, arbitrary precision
# SCALE: self-calibrated anomaly score for flagging atypical / error sequences
cal = graph.calibrate_scale(seed=42) # calibrate once against the graph
print(f"SCALE = {graph.scale_score('CASSLEPSGGTDTQYF', cal):.2f}") # higher = more anomalousFlashBackGraph.from_file(path) accepts any of the formats in Input format above (plain, seq<TAB>count, AIRR-style tabular, FASTA, FASTQ, gzip/bzip2/xz-compressed) and is equivalent to running lzg flashback build path -o out.lzg and then loading the result: same detection, same validation, same graph. LZGraph.from_file(path, variant='aap') is the same guarantee for LZGraph, equivalent to lzg build path -o out.lzg.
from LZGraphs import FlashBackGraph
# Write a tiny example file (one CDR3 per line, or seq<TAB>count for abundance)
with open('repertoire.tsv', 'w') as f:
for s in ['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', 'CASSLGQGSTEAFF']:
f.write(s + '\n')
graph = FlashBackGraph.from_file('repertoire.tsv')
print(graph.n_nodes, 'nodes')For a clean, uncompressed file like this one, from_file streams straight into the C builder in constant memory, so it is the right choice for very large repertoires. For incremental / checkpointed builds over repertoires too large to write to a single file up front, use FlashBackStream: same accumulator with add_sequences(), snapshot(), and finalize(). See the class docstring (help(FlashBackStream)) for the streaming protocol.
Everything above is also available from the terminal under lzg flashback. Input is detected by content, so the FASTA below needs no flags; see Input format for the full list.
# Build a FlashBackGraph. Accepts FASTA, FASTQ, AIRR TSV/CSV, plain,
# seq<TAB>count, optionally gzip/bzip2/xz compressed.
lzg flashback build repertoire.fa -o rep.fb.lzg
# Inspect it
lzg flashback info rep.fb.lzg
# Simulate sequences from the Markovian distribution
lzg flashback simulate rep.fb.lzg -n 10000 --seed 42 > simulated.txt
# Exact generation probability per sequence: TSV of `sequence pgen`
# (natural log) on stdout, one row per query.
lzg flashback score rep.fb.lzg queries.txt
# Exact diversity, and the SCALE anomaly score
lzg flashback diversity rep.fb.lzg
lzg flashback scale rep.fb.lzg queries.txtlzg flashback build and FlashBackGraph.from_file read through the identical pipeline, so the graph is the same either way. Only presentation goes to stderr; the data on stdout is clean, so lzg flashback simulate rep.fb.lzg -n 1000 | head -3 behaves as you would expect in a pipeline.
pgen answers "how probable is this sequence". pseq_analysis() answers the distributional question: across the whole space of sequences the graph can generate, how is probability spread out? It is computed from the power-sum transform M(q) = sum_s P(s)^q as a forward dynamic program over the edges, so the transform, its derivatives, and the surprisal moments that follow are exact rather than estimated from sampled walks.
from LZGraphs import FlashBackGraph
fb = FlashBackGraph(['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', 'CASSLEPQTFTDTFFF',
'CASSLGQGSTEAFF', 'CASSLGIRRT'])
psq = fb.pseq_analysis()
# Exact surprisal moments of the generated distribution
m = psq.moments()
print(f"mean surprisal = {m['mean']:.3f} nats, sd = {m['std']:.3f}")
# Where does one sequence sit in that spectrum?
pos = psq.position('CASSLGIRRT')
print(f"P(seq) = {pos['pseq']:.3g}, surprisal = {pos['surprisal']:.3f}")
print(f"{pos['fraction_of_sequences_at_least_as_probable']:.1%} of sequences are at least as probable")
# Expected distinct sequences after sampling n times
print(psq.expected_richness(10_000)['expected_richness'])Also available on the returned object: mellin(q) / log_mellin(q) / derivatives(q, order) for the transform itself, cumulants(), length_profile() and length_derivatives() for length-stratified mass and moments, exact_atoms() for exhaustive enumeration on small supports, histogram() for a deterministic surprisal-grid reconstruction on large ones (reporting its grid spacing and a rounding-error bound), saddlepoint() for smooth Lugannani-Rice PDF/CDF inversion, and expected_frequency_spectrum(n, max_count).
This is a Python-only API; there is no lzg flashback subcommand for the spectrum. For per-sequence probabilities from the terminal, use lzg flashback score above.
LZGraph |
FlashBackGraph |
|
|---|---|---|
| Tokenization | LZ76 dictionary | FlashBack (run-peeling) |
| Structure | LZ-constrained walks | Markovian DAG |
| Diversity / entropy / path count | Analytical (with MC where needed) | Closed-form forward DP |
| Self-calibrated anomaly scoring (SCALE) | No | Yes |
| V/J gene annotation & gene-conditioned simulation | Yes | No |
| Encoding variants | aap, ndp, naive |
Single representation |
| Sampling-free p-sequence spectrum | No | Yes (pseq_analysis()) |
CLI tool (lzg) |
Yes (lzg build, ...) |
Yes (lzg flashback build, ...) |
| Streaming / incremental build | No | Yes (FlashBackStream) |
Benchmark figures below are from a single CPU core on a 5,000-sequence amino-acid CDR3 repertoire (mean length 14.7 aa; resulting LZGraph has ~1,700 nodes, ~9,600 edges). See docs/resources/benchmarks.md for the full table and methodology.
| Operation | Throughput |
|---|---|
| Graph construction | ~50,000 sequences/sec (5k seqs in <100 ms) |
pgen() scoring |
~5,000 sequences/sec (constant across batch sizes) |
simulate() |
~4,800 sequences/sec |
| Hill numbers via MC (10k walks) | ~2 sec |
Load / save .lzg |
~100× faster than rebuilding |
For repertoires of ~100k sequences and above, graph construction stays linear and saved .lzg files round-trip in seconds. For a clean, uncompressed input file, FlashBackGraph's from_file and FlashBackStream paths operate in bounded memory; we have built and validated graphs with >70,000 nodes and >11M edges this way. (Compressed or headered input is decompressed and validated in memory first, same as lzg build; see Input format.)
Every snippet in this section is paste-and-runnable after the Setup block below. graph flags a method that works on either class; lz_graph is an LZGraph instance and fb_graph is a FlashBackGraph instance. Methods marked LZGraph-only or FlashBackGraph-only are not implemented on the other class.
from LZGraphs import LZGraph, FlashBackGraph, jensen_shannon_divergence
seqs = ['CASSLEPSGGTDTQYF', 'CASSDTSGGTDTQYF', 'CASSLEPQTFTDTFFF',
'CASSLGQGSTEAFF', 'CASSLGIRRT']
v_genes = ['TRBV5-1*01', 'TRBV5-1*01', 'TRBV5-1*01', 'TRBV7-2*03', 'TRBV7-2*03']
j_genes = ['TRBJ2-7*01', 'TRBJ2-7*01', 'TRBJ2-7*01', 'TRBJ1-2*01', 'TRBJ1-2*01']
lz_graph = LZGraph(seqs, variant='aap', v_genes=v_genes, j_genes=j_genes)
fb_graph = FlashBackGraph(seqs)
graph = lz_graph # `graph` flags methods that work on either class
graph_a = LZGraph(seqs[:3], variant='aap')
graph_b = LZGraph(seqs[2:], variant='aap')
population = LZGraph(seqs * 4, variant='aap')
patient_seqs = ['CASSLGIRRT', 'CASSLGQGSTEAFF']
lz_reference = population
lz_sample = LZGraph(seqs, variant='aap')# Log-probability of a sequence (works on LZGraph and FlashBackGraph alike)
graph.pgen('CASSLEPSGGTDTQYF') # single → float
graph.pgen(seqs[:3]) # batch → np.ndarray
# Simulate (both classes)
result = graph.simulate(1000, seed=42)
result = lz_graph.simulate(100, v_gene='TRBV5-1*01', j_gene='TRBJ2-7*01') # LZGraph onlygraph.effective_diversity() # exp(Shannon entropy)
graph.hill_number(2) # inverse Simpson
graph.hill_numbers([0, 1, 2, 5]) # multiple orders → np.ndarray
# Both classes
graph.path_count # distinct walks (exact int on FlashBackGraph)
graph.pgen_moments() # moments of the log-pgen distribution
graph.pgen_distribution() # Gaussian-mixture approximation of log-pgen
# LZGraph-only
lz_graph.predicted_richness(100_000) # expected unique seqs at depth
lz_graph.predicted_overlap(10000, 50000) # expected shared sequences
lz_graph.predict_sharing([1000]*5, max_k=5) # sharing spectrum across donors
# FlashBackGraph-only (closed-form)
fb_graph.pseq_analysis() # sampling-free p-sequence spectrum
cal = fb_graph.calibrate_scale(seed=0) # self-calibrate the SCALE anomaly score (once)
fb_graph.scale_score('CASSLEPSGGTDTQYF', cal) # SCALE: higher = more anomalouscombined = graph_a | graph_b # union (LZGraph and FlashBackGraph)
shared = graph_a & graph_b # intersection (both)
unique_a = graph_a - graph_b # difference (both)
personal = population.posterior(patient_seqs, kappa=10.0) # Bayesian update (both)jsd = jensen_shannon_divergence(graph_a, graph_b) # natural log (nats): 0.0 identical, ln(2) ≈ 0.693 disjointgraph.feature_stats() # 15-element summary vector (both classes)
graph.feature_mass_profile() # position-based mass distribution (both classes)
# LZGraph-only
lz_reference.feature_aligned(lz_sample) # project sample into a fixed reference space# Both classes use the same .lzg binary format, but each file is class-specific.
lz_graph.save('rep_lz.lzg')
loaded_lz = LZGraph.load('rep_lz.lzg')
fb_graph.save('rep_fb.lzg')
loaded_fb = FlashBackGraph.load('rep_fb.lzg')Full documentation with tutorials, concept guides, and API reference:
https://MuteJester.github.io/LZGraphs/
- Quick Start: build your first graph in 5 minutes
- Tutorials: graph construction, sequence analysis, diversity metrics
- API Reference: complete class and function reference
- CLI Reference: terminal tool documentation
If you use LZGraphs in published research, please cite the methods paper. If you also want to cite a specific software version, add the software entry below.
@article{konstantinovsky2023novel,
title={A novel approach to T-cell receptor beta chain ({TCRB}) repertoire encoding using lossless string compression},
author={Konstantinovsky, Thomas and Yaari, Gur},
journal={Bioinformatics},
volume={39},
number={7},
pages={btad426},
year={2023},
publisher={Oxford University Press},
doi={10.1093/bioinformatics/btad426}
}
@software{lzgraphs_software,
author={Konstantinovsky, Thomas},
title={{LZGraphs}: {LZ76} and {FlashBack} compression graphs for immune repertoire analysis},
url={https://github.com/MuteJester/LZGraphs},
year={2026}
}Contributions are welcome. Please open an issue or submit a pull request.
LZGraphs builds a CPython extension from a C library at install time, so a working C toolchain is required:
- Linux:
gccorclang(any version supporting C11) - macOS: Xcode command-line tools (
xcode-select --install) - Windows: Visual Studio Build Tools with the "Desktop development with C++" workload
Then:
git clone https://github.com/MuteJester/LZGraphs.git
cd LZGraphs
pip install -e ".[dev]" # editable install + dev extras (pytest, pytest-cov, ruff, scipy, build)
pytest # run the test suite (~1,700 tests)- Fork the repository
- Create a feature branch (
git checkout -b feature/my-feature) - Add tests for new functionality; make sure
pytestandpytest tests/regression/both pass - Commit your changes (small, focused commits preferred)
- Push and open a Pull Request describing the motivation and any API changes
MIT License. See LICENSE for details.
Thomas Konstantinovsky, thomaskon90@gmail.com
GitHub · PyPI · Documentation