Abstract
Rice disease detection systems based on visual deep learning achieve high classification accuracy but fail to enforce biological plausibility, provide pre symptom risk forecasting or deliver agronomically interpretable explanations. This paper presents RiceMultiModalNet, a proof-of-concept three-branch late-fusion framework. Evaluated on a computationally synthesised from independent public repository multimodal dataset in which leaf images, environmental variables, and phenological stage labels were not co-collected from the same field locations or time points. Results demonstrate the feasibility of three operational capabilities under controlled conditions and should not be interpreted as validated field-deployment or real-world early-warning capability. It integrates a ResNet-50 visual classifier, a phenological growth stage encoder, and an eight variable meteorological risk (multilayer perceptron) MLP in a single pipeline, evaluated on a computationally synthesised dataset of 4804 rice leaf images from publicly available repositories (Mendeley, Kaggle) together with NASAPOWER reanalysis environmental data assigned computationally. Three operational contributions are reported as (1) an 8 × 3 literature derived phenological susceptibility constraint module suppressing biologically implausible prediction at inference time, (2) an image free environmental risk sub model enabling pre symptom disease advisory from meteorological inputs alone, and (3) SHAP-based agronomic explainability linking predictions to phenological and environmental drivers, spatially validated by Grad-CAM (mean lesion energy concentration 80.2%, mean IoU 46.6, pointing game accuracy 60–65%, full per disease breakdown in Table 11). The framework’s primary contributions are three operational capabilities (1) an 8 × 3 phenological susceptibility constraint module suppressing biologically implausible predictions, (2) an image-free environmental risk sub-model (AUC = 0.99%, ECE = 0.0281) enabling pre-sympton disease advisory from meterological inputs alone, and (3) SHAP based agronomic explainability validated spatially by Grad-CAM (mean lesion energy concentration 80.2%, mean IoU 46.6%, pointing game accuracy (60–65%). The visual branch achieved 99.58% classification accuracy on the benchmark, confirming correct visual branch functionality, however all ablation configurations including the single-model visual baseline achieved identical accuracy, confirming that multi-modal value resides entirely in the three operational capabilities described above. All three modalities were computationally synthesised rather than co-collected in the field. The AUC and SHAP rankings may reflect algorithmic correlations introduced during synthesis and should be treated as proof-of- concept indicators only. Longitudinal field validation with co-collected multimodal data remains an essential next step.
Acknowledgements
The authors sincerely acknowledge the use of AI tools, which contributed to editing and improving the readability of this article. However, the content, ideas, and conclusion presented are entirely the work of the authors, reflecting their original research and insights.
Funding
Open access funding provided by Vellore Institute of Technology.
Author information
Authors and Affiliations
Corresponding author
Ethics declarations
Competing interests
The authors declare no competing interests.
Additional information
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Vijayan, S., Chowdhary, C.L. RiceMultiModalNet: a proof-of-concept biologically constrained multi-modal deep learning for rice disease detection integrating phenological stage and environmental risk.
Sci Rep (2026). https://doi.org/10.1038/s41598-026-65961-z
Received:
Accepted:
Published:
DOI: https://doi.org/10.1038/s41598-026-65961-z
Keywords
- Environmental modelling
- Explainable artificial intelligence (XAI)
- Phenological stages
- Rice disease detection
- Sustainable agriculture
Source: Ecology - nature.com
