Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study

AI in healthcare
Published: arXiv: 2608.27578v1
Authors

Nathaniel Chen Kouroche Bouchiat Peter Steiner Azarakhsh Jalalvand SangKyeun Kim Egemen Kolemen

Abstract

Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such a model: with over 20 sensor types spanning 5 orders of magnitude in sampling rate, mixed tensor structures (point measurements, spectrograms, images), and nonstationary physics. We analyze our input complexity and discuss trade-offs between temporal context and frequency resolution. Our analysis provides a template for representing multi-modal fluctuation data at scale, with implications for both multi-modal control systems and nuclear fusion.

Paper Summary

General Summary
This paper addresses important challenges in the field by introducing novel methodologies and approaches. The research contributes to advancing the state-of-the-art and has practical applications that could benefit various domains.
Paper Information
Categories:
physics.plasm-ph cs.LG
Published Date:

arXiv ID:

2608.27578v1

Quick Actions