Trustworthy decision support systems utilizing a multimodal approach (MMA) integrate diverse data modalities to enhance robustness, transparency, and fairness in artificial intelligence (AI) applications. In this study, we present an MMA for decision support in the Air Traffic Management (ATM) domain, particularly within Remote Digital Towers (RDTs). RDTs replace traditional control towers with AI-driven digital solutions, enhancing operational efficiency. Our approach addresses key multimodal challenges—translation, alignment, and co-learning—by implementing (a) an open-vocabulary-based object detection model for video processing and (b) an audio-to-text transcription and semantic word identification model. The YOLO-World deep-learning model is employed for object detection, while audio data analysis takes advantage of a benchmark data set, semantic identification techniques, and explainability. Additionally, the system integrates robust machine learning techniques, including data augmentation and perturbation, to maintain consistent performance across varied operational conditions. This proof-of-concept demonstrates the potential of multimodal AI systems to enhance decision support and improve safety in ATM environments.