Speaker
Description
Over the past 25 years, over 2 million X-ray sources have been serendipitously discovered by various X-ray observatories, where the majority remain unclassified. Traditional manual classification methods alone are increasingly unable to keep up with the growth of data. We present the results and lessons learned from applying a random forest classifier to X-ray sources from the XMM-Newton observatory, demonstrating how machine learning can augment traditional classification methods by rapidly classifying common source classes (i.e., stars and active galactic nuclei) and identifying rare source class candidates (i.e., neutron stars and black holes) for manual review. A primary challenge is the severe class imbalance in the X-ray source population, combined with missing values due to non-detection of optical and infrared counterparts, which provide crucial information on the nature of X-ray sources. We employed a physically motivated oversampler alongside random SMOTE and treated missing values as informative features. Our result shows that while the random forest algorithm achieves high accuracy for common source classes, performance degrades substantially for rare classes, highlighting the effectiveness of machine learning as a filter for common source classes and the necessity of more effective oversampling methods to improve the accuracy of rare source classification.
| Do you plan to submit a 4-page extended abstract on OpenReview (only for Presentations/Posters)? | No |
|---|