Machine Learning and Data Engineering
Our research connects machine learning and data engineering to make AI models useful under real-world constraints on resources, data, and computing systems. We design learning methods by considering data access, computing architecture, and available resources together from the outset. This perspective spans efficient access to very large data collections, scalable execution on high-performance systems, and compact models for edge and TinyML hardware. Current funded projects include AI4Forest and TinyAIoT.
Large-Scale Machine Learning
Many large-scale applications are limited less by a model's arithmetic than by data access and movement. We therefore couple learning methods with index structures, storage systems, and execution pipelines.
In search-by-classification, a small number of examples describe the objects a user wants to retrieve. Our index-aware decision trees translate inference into range queries, enabling interactive searches across billions of objects. RapidEarth applies this idea to satellite-image archives, while CLIP-Branches combines it with interactive fine-tuning of multimodal models.
- Lülf et al. (2024). CLIP-Branches: Interactive Fine-Tuning for Text-Image Retrieval. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Demo Track. SIGIR 2024.
- Lülf et al. (2023). RapidEarth: A Search Engine for Large-Scale Geospatial Imagery. Proceedings of the 31st International Conference on Advances in Geographic Information Systems, Demo Paper. SIGSPATIAL 2023.
- Lülf et al. (2023). Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests. Proceedings of the VLDB Endowment, 16, 2845–2857. VLDB 2023.
Current research projects often process hundreds of terabytes of satellite observations. Here, transferring the data can become as important a bottleneck as computation. We develop trainable selection masks that identify and transfer only the parts of an input that matter for a task. For petabyte-scale archives, automated preselection can substantially reduce data movement and end-to-end inference time.
This line of work directly supports the scalable Earth-observation pipelines developed in AI4Forest.
- Oehmcke & Gieseke (2022). Input Selection for Bandwidth-Limited Neural Network Inference. Proceedings of the 2022 SIAM International Conference on Data Mining, 280–288. SDM 2022.
- Pauls et al. (2024). Estimating Canopy Height at Scale. 41st International Conference on Machine Learning. ICML 2024.
Tiny Machine Learning
TinyML brings training and inference to severely resource-constrained devices. Adapted training procedures and memory layouts reduce the resource requirements of boosted-tree models by factors of 4–16 while preserving predictive performance. Trainable quantization further reduces the sensor data that must be transmitted. Applications include privacy-preserving bicycle counting and energy-efficient bird-species recognition.
These methods are developed and evaluated in the TinyAIoT project and related environmental-monitoring initiatives such as Birdiary.
- Herrmann et al. (2026). Boosted Trees on a Diet: Compact Models for Resource-Constrained Devices. The Fourteenth International Conference on Learning Representations. ICLR 2026.
- Schrödter et al. (2026). Trainable Bitwise Soft Quantization for Input Feature Compression. Third Conference on Parsimony and Learning. CPAL 2026.
- Kurkela et al. (2026). TinyML for Environmental Monitoring: Bird Species Image Classification on Resource-Constrained Devices. European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Applied Data Science Track. ECML PKDD 2026.
Machine Learning and High-Performance Computing
Our work exploits the parallelism of modern many-core systems through hardware-aware algorithm design. Buffer k-d trees batch search requests for massively parallel nearest-neighbor search on GPUs. Instead of traversing a tree independently for every query, buffers collect queries at tree nodes and process them together. This reorganizes irregular control flow into large, hardware-friendly batches and improves memory access on GPUs.
We have extended this perspective to data-intensive learning models and scientific pipelines, including parallel regression for satellite time series. In change detection, such methods can reduce computations over billions of time series from weeks or years to hours or days. The central principle is to co-design algorithms, memory movement, vectorization, and distributed execution for the target architecture.
- Gieseke et al. (2020). Massively-Parallel Change Detection for Satellite Time Series Data with Missing Values. Proceedings of the 36th IEEE International Conference on Data Engineering, 385–396. ICDE 2020.
- Gieseke & Igel (2018). Training Big Random Forests with Little Resources. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1445–1454. KDD 2018.
- Gieseke et al. (2014). Buffer k-d Trees: Processing Massive Nearest Neighbor Queries on GPUs. Proceedings of the 31st International Conference on Machine Learning, 172–180. ICML 2014.
Applications
Alongside our methodological research, we develop machine-learning techniques with experts from Earth observation, astrophysics, and other application domains.
Earth Observation
Our work ranges from detecting abrupt ecosystem changes and mapping individual trees and their carbon stocks to creating global, time-dependent canopy-height maps. We build efficient inference pipelines for petabyte-scale collections, develop foundation representations for environmental monitoring, and study predictive uncertainty through quantile regression.
Much of this work is carried out within AI4Forest; interactive results and further material are also available at ai4forest.eu.
- Fayad et al. (2025). DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications. 42nd International Conference on Machine Learning. ICML 2025.
- Pauls et al. (2025). Capturing Temporal Dynamics in Large-Scale Canopy Tree Height Estimation. 42nd International Conference on Machine Learning. ICML 2025.
- Schrödter et al. (2026). Canopy Tree Height Estimation using Quantile Regression: Modeling and Evaluating Uncertainty in Remote Sensing. Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics. AISTATS 2026.
- Bernardino et al. (2025). Predictability of Abrupt Shifts in Dryland Ecosystem Functioning. Nature Climate Change, 15, 86–91.
Astrophysics
Astronomy and particle physics were important application areas in an earlier phase of our research. Modern sky surveys produce very large image and catalogue collections in which rare, scientifically relevant objects must be identified among millions or billions of observations. Manual inspection is therefore impossible, making accurate and scalable machine-learning pipelines essential.
For astronomical surveys, we developed nearest-neighbor methods for photometric redshift estimation and for discovering previously unknown high-redshift quasars. We also designed convolutional neural networks that distinguish genuine transient events from imaging artefacts and thereby reduce the number of candidates requiring expert review. These projects established our focus on scalable inference, efficient candidate selection, and close collaboration with domain scientists.
- Gieseke et al. (2017). Convolutional Neural Networks for Transient Candidate Vetting in Large-Scale Surveys. Monthly Notices of the Royal Astronomical Society, 472, 3101–3114. MNRAS.
- Polsterer, Zinn & Gieseke (2013). Finding New High-Redshift Quasars by Asking the Neighbours. Monthly Notices of the Royal Astronomical Society, 428, 226–235. MNRAS.