Reliable Retrieval-Augmented Feature Generation with Large Language Model Reasoning

Published in Knowledge and Information Systems, Volume 68, Article 172, 2026

Feature generation can improve learning in data-limited domains, but useful and interpretable features often require specialized knowledge. This paper introduces Retrieval-Augmented Feature Generation (RAFG), a training-free framework that retrieves domain knowledge based on relationships among existing features and uses large language model reasoning to synthesize new candidates. Reasoning-aware re-ranking, causal alignment checks, and counterfactual validation help distinguish useful features from unsupported or noisy generations. Experiments across medical, economic, and geographic datasets show that RAFG produces meaningful features and improves domain-specific classification performance over competitive baselines.

📄 In Knowledge and Information Systems, Volume 68, Article 172.
🔗 Paper Link (Springer)