Microsoft Research has released Skala 1.1, an updated deep-learning exchange-correlation functional for density functional theory (DFT). The new version is trained on 2.5 times more reference data than its predecessor.
The expanded training set draws from the Microsoft Research Accurate Chemistry Collection, which now includes electron affinities and noncovalent clusters alongside existing thermochemistry and kinetics data.
On the GMTKN55 benchmark suite, Skala 1.1 achieves a weighted average error of 2.8 kcal/mol across 55 categories. It earns top marks in 32 of those categories, spanning thermochemistry, reaction kinetics, and non-covalent interactions.
At the computational cost of a meta-GGA functional, Skala 1.1 outperforms the best global hybrid functionals while retaining semi-local efficiency.
What's new
- Skala 1.1: trained on 2.5× more data from the expanded MSR-ACC, including new categories such as electron affinities and noncovalent clusters
- Benchmark performance: 2.8 kcal/mol weighted average error on GMTKN55; gold medals in 32 of 55 categories
- Software integrations: available now in CP2K; integrations in progress for Psi4, FHI-aims, ORCA, and VASP
- Living benchmark: continuous performance tracking across codes and hardware to accelerate optimization
- Access channels: Skala Community Edition on GitHub (GPU4PySCF + ASE), Microsoft Foundry Labs, and native CP2K integration
Why it matters
DFT powers molecular simulation across chemistry, materials science, catalysis, energy technologies, and drug discovery. The field has long faced a trade-off: high-accuracy hybrid functionals are computationally expensive, while efficient semi-local functionals sacrifice predictive reliability.
Skala's deep-learning approach aims to break this trade-off by learning exchange-correlation functionals directly from massive wavefunction-quality datasets.
The CP2K integration is significant because CP2K excels at large-scale molecular dynamics. Combining CP2K's scalability with Skala's accuracy could enable predictive simulations of complex systems — such as catalytic interfaces or battery materials — at previously inaccessible time and length scales.
The living benchmark addresses a practical adoption barrier: researchers need confidence that a new functional performs consistently across different codes and hardware before committing computational resources.
Our take
The continuous-improvement philosophy — where each Skala release supersedes the last rather than joining a "functional zoo" — is a pragmatic shift. However, the 2.8 kcal/mol average error still sits above the 1 kcal/mol chemical accuracy threshold needed for truly predictive modeling without empirical correction. The real test is whether the living benchmark drives optimization that closes this gap while maintaining semi-local speed across the expanding ecosystem of supported codes.