Skip to content

About

Code for learning lattice parameters from powder XRD patterns via an invariant lattice bispectrum representation.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

bispectrum-xrd-ml

Predicting unit cells from powder XRD data: a transformer maps simulated/experimental XRD patterns to a crystallographic bispectrum, which is then inverted (L-BFGS) to recover lattice parameters. See training/README.md for the full pipeline, model configs, and training/eval commands.

This is the code for Learning Lattice Parameters from Powder X-Ray Diffraction Data Using Invariants (Hofgard, Min, Segal, Mittan-Moreau, Mansouri Tehrani, Oklejas, Nigam, Paley, Brewster, Smidt), arXiv:2607.21829. See Citation.

Structure

  • utilities.py — shared core library (XRD simulation, bispectrum computation, reciprocal-lattice neighbor search, Niggli/primitive-lattice comparisons). Used by both training/ and bispectrum/.
  • training/ — transformer training and evaluation (Materials Project data, RRUFF and opXRD CNRS experimental patterns), plus mp_full/ data augmentation (strain, texture, Caglioti, background noise) and the peak-list input variant. Covers both the mp20 and mpfull datasets.
  • bispectrum/ — L-BFGS inversion of a predicted bispectrum back to lattice parameters (run_alg_inversion.py); the lookup-database builder (build_lookup_database.py) and its cost measurements (time_bispec_database.py); the cctbx-based comparison of predicted and true lattices (inversion_results_cctbx.py) and the plots built on it; inversion timing; and the XRD-pattern leakage audit (mp20_xrd_leakage.py, mpfull_xrd_leakage.py, plot_similarity_vs_error.py). See bispectrum/README.md.
  • tutorials/ — example notebooks. Some fetch structures from the Materials Project via mp-api (see Installation below).

Installation

pip install -e .            # core package (utilities.py + runtime deps)
pip install -e ".[dev]"     # + pytest
pip install -e ".[notebooks]"  # + mp-api, only needed for the tutorial notebooks

training/ and bispectrum/ scripts are not part of the installed package (they're flat scripts that import each other and utilities directly); run them from within their own directory, e.g. cd training && python train.py ....

Known issue: mp-api's dependency chain (via emmet-core) can conflict with very recent pymatgen releases and force a downgrade if installed carelessly. Install the notebooks extra into a scratch/throwaway environment first if you need it, rather than a shared one.

Environment variables

variable used by meaning
POWDERXRD_DATA_ROOT all scripts and hydra configs directory holding your datasets, checkpoints and results (required; there is no default)
POWDERXRD_REPO_DIR the SLURM scripts (training/*.sh, bispectrum/*.sh) absolute path of this checkout; a batch job cannot locate it itself
CNRS_JSON_PATH training/prep_opxrd_cnrs.py --source aggregated optional; path of an aggregated CNRS JSON (or pass --json-path, or use --source raw)
POWDERXRD_PYTHON training/run_rruff_inversion.sh optional; Python interpreter for its inline helper (default python3)

The SLURM scripts contain #SBATCH -A YOUR_ACCOUNT and, in the hydra launcher config, slurm_account: YOUR_ACCOUNT; replace these with your own project before submitting.

Data

No datasets, trained checkpoints or inversion results are included in this repository. All training, evaluation and inversion scripts read from the data root above (the subdirectory layout is described in training/README.md):

export POWDERXRD_DATA_ROOT=/path/to/your/data

The data used in the paper comes from:

  • Materials Project structures (full set and the MP-20 subset), from which the simulated patterns and bispectra are generated by the scripts in training/mp_full/. The Materials Project ids of every split are listed in training/splits/: mpfull_{train,val,test}_ids.txt (92,708 / 31,722 / 30,449; split by reduced formula, 60/20/20, seed 42, by training/mp_full/data_gen_full.py) and mp20_{train,val,test}_ids.txt (27,136 / 9,047 / 9,046; the standard MP-20 splits of CDVAE, as used by Crystalyze).
  • RRUFF experimental powder patterns and structures, from the RRUFF project (https://www.rruff.net/). They are not redistributed here: the entries used are listed in training/splits/rruff_alpha_ids.txt (228 entries) and training/splits/rruff_crystalyze_ids.txt (148 entries), and the patterns and structures must be downloaded from RRUFF. The reported Crystalyze RRUFF results use the 134 entries in training/splits/rruff_crystalyze_eval_ids.txt: of the 148, 7 overlap the MP-20 training set and 7 have simulated single-crystal powder profiles rather than measured powder patterns, and those 14 are excluded. The processed training/RRUFF/xrd_data.pkl is not tracked either, because of its size. If you use RRUFF data, please cite: Lafuente, B., Downs, R. T., Yang, H., & Stone, N. (2015). The power of databases: the RRUFF project. In Highlights in Mineralogical Crystallography, T. Armbruster and R. M. Danisi, Eds., Berlin, Germany, W. De Gruyter, 1–30.
  • opXRD (open experimental powder XRD database), CNRS/COD contribution, used for the real-data evaluation in training/prep_opxrd_cnrs.py and the *cnrs* scripts. The ids of the independently quality-controlled subset used for evaluation (599 patterns) are listed in training/splits/cnrs_verified_ids.txt; the pattern data itself is not included. The real-data workflow is described in training/README.md.

The lookup database used by bispectrum/run_alg_inversion.py (-mat_proj_df) is built by bispectrum/build_lookup_database.py. It holds the bispectrum and reciprocal lattice of every Materials Project structure in the full set (train, validation and test splits together), from the output of training/mp_full/data_gen_full.py --mode bispec; at inference the query material itself is excluded from the search. The database used for the paper additionally contains the 1,922 MP-20 structures that are not in the full-MP splits (156,801 entries in total); the script reproduces the other 154,879 entries exactly.

Large notebooks with saved outputs and the MP-20 CSV files under training/data/ were removed from version control to keep the repository small; the MP-20 split ids are in training/splits/.

Testing

pytest tests/

Tests cover the utilities.py functions actually used by training/ and bispectrum/ (XRD simulation, bispectrum computation and its rotation invariance, reciprocal-lattice neighbor search, Niggli/primitive lattice comparisons). A block of utilities.py (Selling reduction and a few other lattice-geometry helpers) has no callers anywhere in the repo and is deliberately left untested rather than removed.

Citation

@article{hofgard2026lattice,
  title   = {Learning Lattice Parameters from Powder X-Ray Diffraction Data Using Invariants},
  author  = {Hofgard, Elyssa and Min, Kyucheol and Segal, Nofit and Mittan-Moreau, David W. and
             Mansouri Tehrani, Aria and Oklejas, Vanessa and Nigam, Jigyasa and Paley, Daniel W. and
             Brewster, Aaron S. and Smidt, Tess},
  year    = {2026},
  eprint  = {2607.21829},
  archivePrefix = {arXiv}
}

About

Code for learning lattice parameters from powder XRD patterns via an invariant lattice bispectrum representation.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages