This repository provides the minimum dataset required to verify and reproduce the results reported in the following manuscript:
@unpublished{mao2026neurosymbolic,
title = {Neurosymbolic AI: Fundamental Learning for Context Engineering in Computational Linguistics},
author = {Mao, Rui and Huang, Zihao and Zhang, Xulang and Cambria, Erik},
journal = {npj Artificial Intelligence},
note = {Manuscript under review},
year = {2026}
}The evaluation data included in this repository were derived from the following publicly available sources:
Tiuleneva, M., Porvatov, V. A., and Strapparava, C. Big-Five Backstage: A Dramatic Dataset for Characters Personality Traits & Gender Analysis. In Proceedings of the Workshop on Cognitive Aspects of the Lexicon @ LREC-COLING 2024, pp. 114–119, 2024.
Misra, R., and Arora, P. Sarcasm detection using news headlines dataset. AI Open, 4:13–18, 2023. https://doi.org/10.1016/j.aiopen.2023.01.001
Mao, R., He, K., Ong, C. B., Liu, Q., and Cambria, E. MetaPro 2.0: Computational metaphor processing on the effectiveness of anomalous language modeling. In Findings of the Association for Computational Linguistics: ACL, pp. 9891–9908, 2024.
Chen, Z., and Qian, T. Relation-aware collaborative learning for unified aspect-based sentiment analysis. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 3685–3694, 2020.
Go, A., Bhayani, R., and Huang, L. Twitter sentiment classification using distant supervision. CS224N Project Report, Stanford, 1(12), 2009.
The training data used in this study were derived from the following publicly available sources:
Xu, W., Ritter, A., Baldwin, T., and Rahimi, A. (eds.). Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021). 2021.
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2):313–330, 1993.
Sang, E. F., and Buchholz, S. Introduction to the CoNLL-2000 shared task: Chunking. arXiv preprint cs/0009008, 2000.
Miller, G. A., Leacock, C., Tengi, R., and Bunker, R. T. A semantic concordance. In Human Language Technology: Proceedings of a Workshop Held at Plainsboro, New Jersey, 1993.
Hulth, A. Improved automatic keyword extraction given more linguistic knowledge. In Proceedings of the 2003 Conference on Empirical Methods in Natural Language Processing, pp. 216–223, 2003.
Tjong Kim Sang, E. F., and De Meulder, F. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pp. 142–147, 2003.
Pradhan, S., Moschitti, A., Xue, N., Uryupina, O., and Zhang, Y. CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes. In Joint Conference on EMNLP and CoNLL – Shared Task, pp. 1–40, 2012.
Pang, B., and Lee, L. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058, 2004.
- This repository contains processed data intended for reproducibility and verification purposes.
- The original datasets remain the property of their respective authors and providers.
- Please consult the original sources for complete dataset descriptions, licensing terms, and access conditions.