pyfracval.catalog#
Load and filter a cluster_index.csv master index of generated aggregates.
cluster_index.csv is written by scripts/generate_cluster_data.py as a
flat table – one row per generated .dat file, with the generation
parameters (sigma/Df/N/kf/…) that would otherwise only be recoverable by
parsing directory names or re-reading every file’s YAML header via
pyfracval.schemas.Metadata.from_file(). This module is the native way to
query that table instead: no pandas dependency (only an optional/plot
dependency of this package, not a core one), just the stdlib csv module
and a small pydantic model (pyfracval.schemas.ClusterEntry).
Module Contents#
- pyfracval.catalog.load_catalog(index_path, data_root=None)[source]#
Load every row of a
cluster_index.csvas a list ofClusterEntry.index_path’s
filepathcolumn holds an absolute path from wherever the data was originally generated – not portable across machines or repos that keep their own copy of the same aggregate files. When data_root is given, each entry’sfilepathis re-based under it instead of trusting the stored absolute path: the last 3 path components (<config>/<sigma_..__Df_..__N_..>/<filename>.dat) are the part of the layoutscripts/generate_cluster_data.pyactually guarantees, so re-joining just those under data_root is reliable regardless of where the CSV says the file used to live. When data_root isNone,filepathis used as-is (verbatim from the CSV).
- pyfracval.catalog.filter_catalog(entries, *, sigma=None, Df=None, N=None, kf=None, config=None, success_only=True)[source]#
Filter catalog entries by generation parameters.
Each keyword accepts either a single value (exact match) or an iterable of allowed values;
None(the default for all but success_only) means “don’t filter on this field”. success_only (defaultTrue) drops any entry withsuccess=Falsebefore applying the other filters – a failed generation attempt has no usable geometry.