RESEARCH

Research Overview

My research focuses on discovering and studying rare, unusual, and overlooked astrophysical systems in the large datasets produced by modern astronomical surveys. I use machine learning — particularly anomaly detection, active learning, and representation learning — to identify objects that would be difficult to find through conventional searches, and then use these discoveries to investigate their underlying astrophysics.

My work spans both the dynamic Universe and galaxy evolution. In time-domain surveys such as ZTF, I study unusual variable sources and explore how information normally discarded as instrumental artefacts can be recovered for astrophysical analysis. In Euclid imaging, I am developing searches for rare galaxy populations and unusual morphologies, including interacting systems, ring galaxies, and other objects that can provide insight into galaxy interactions and evolution.

A central idea behind my research is that machine learning should be more than a classification tool. I use it as a discovery mechanism: to construct new astrophysical samples, identify unexpected phenomena, and reveal regions of parameter space that predefined selections can miss. The next step is then the astronomy — characterising these systems, understanding their physical properties and environments, and determining what they can tell us about the evolution and behaviour of the Universe.

This discovery-to-characterisation approach underpins my long-term research programme for surveys including Euclid and the Vera C. Rubin Observatory, where the scale of the data makes new approaches to scientific discovery essential.

A sample of ring galaxies from the Euclid Q1 galaxy morphology catalogue isolated using active anomaly detection (Sreejith & Nichol 2026).

The Dynamic and Variable Universe

Time-domain surveys are revealing an enormous diversity of variable and transient behaviour, but their scale makes it increasingly difficult to explore objects that fall outside established classes or selection criteria. A major part of my research uses Zwicky Transient Facility (ZTF) light curves to search for unusual variable sources and to investigate populations that may be missed by conventional searches.

I use active anomaly-detection methods to explore large light-curve datasets with an astronomer in the loop. Rather than assigning every source to a predefined class, these methods allow the search to evolve as interesting objects are identified. My work comparing different approaches has shown that their effectiveness can vary substantially depending on the dataset and scientific target, motivating the development of discovery strategies that combine complementary methods rather than relying on a single ranking algorithm.

The objects identified through these searches provide the starting point for astrophysical investigation. I use their light-curve behaviour, periodicity, catalogue information, and available multi-wavelength data to determine the nature of unusual sources and to identify promising targets for further study and follow-up.

This work is also preparation for the Vera C. Rubin Observatory, where the scale and diversity of the time-domain sky will make it impossible to define every scientifically interesting population in advance. My goal is to use adaptive discovery methods to find unusual sources efficiently and, more importantly, turn those discoveries into samples for studying stellar variability, transient phenomena, and unexpected behaviour in the dynamic sky.

Recovering Astrophysics from Survey Artefacts

Examples of the different kinds of artefacts detected using PineForest from 26 fields in ZTF DR3 (Sreejith + 2026).

Astronomical surveys routinely discard detections affected by saturation, instrumental effects, or processing failures. My work explores a different question: can some of these artefacts retain useful information about the astrophysical sources that produced them?

Within the SNAD collaboration, I used active anomaly detection to identify and characterise unusual detections in ZTF, leading to the first large expert-labelled public catalogue of ZTF artefacts. Visual inspection of these objects revealed that the artefacts were not simply a collection of contaminants: some were associated with real variable sources, while others preserved signatures of the variability of bright stars whose direct photometry was compromised by saturation.

I am now investigating these saturated-star echoes as an alternative source of time-domain information. By constructing light curves from multiple artefacts associated with the same bright star, I have shown that their variability can be recovered and that periodic signals can be measured even when conventional photometry of the source is problematic. Recovering consistent periods independently from different echoes of the same star provides an important validation that the variability is astrophysical rather than an artefact of the individual detection.

This work illustrates a broader theme of my research: objects rejected by standard processing pipelines can themselves become scientifically useful datasets. Identifying where astrophysical information survives in these failure modes offers a way to recover sources and measurements that would otherwise be lost.

Example of Mira variable’s period being calculated from a variable echo (Sreejith + 2026).

future direction: Machine Learning for Astronomical Discovery

Finding rare objects in modern astronomical surveys requires methods that can search millions of sources without assuming in advance what an interesting object should look like. I develop and apply active anomaly-detection and representation-learning methods as tools for navigating these datasets and constructing samples for subsequent astrophysical study.

A major part of my recent work has been a systematic comparison of three active anomaly-detection frameworks — PineForest, Astronomaly, and Protégé — across two very different astronomical datasets: ZTF light curves and Euclid galaxy images. The methods show markedly different behaviour between the two domains, demonstrating that there is no universally optimal anomaly detector and that the choice of method can strongly influence the populations that are discovered. I have also investigated how prior expert labels affect subsequent searches, including whether knowledge acquired by one method can be transferred effectively to another.

These experiments motivate my broader approach of using complementary discovery methods rather than treating a single anomaly ranking as definitive. Expert feedback can progressively steer searches towards a scientific question, while unsupervised components retain sensitivity to unexpected objects. For imaging applications, I am also exploring self-supervised representation learning and dimensionality-reduction techniques to describe galaxy morphology without relying entirely on predefined catalogue features or labelled training sets.

The purpose of these methods is ultimately scientific: to reduce enormous survey datasets to manageable, physically interesting samples that can then be characterised in detail. This combination of scalable machine learning and expert-guided exploration provides the methodological foundation for my work on unusual galaxies, variable sources, and other rare phenomena in Euclid, ZTF, and future Rubin data.