Quranic: Extended Quranic Treebank (EQTB) Dataset
| dc.contributor.author | Nashir, Wadee A. | |
| dc.contributor.author | Mohsen, Abdulqader M. | |
| dc.contributor.author | Al-Shargabi, Asma A. | |
| dc.contributor.author | Nour, Mohamed K. | |
| dc.contributor.author | Al-Onazi, Badriyya B. | |
| dc.date.accessioned | 2026-09-03T00:03:17Z | |
| dc.date.issued | 2025-04-25 | |
| dc.description | Institutional repository copy of Version 1 of the dataset originally published as "Quranic" in Mendeley Data (DOI: 10.17632/rk96pn66m4.1). The related Data in Brief article is DOI: 10.1016/j.dib.2025.111940. Third-party source components incorporated into the dataset remain subject to their respective licenses and attribution requirements. | |
| dc.description.abstract | The Extended Quranic Treebank (EQTB) is a computationally accessible, multi-layered linguistic dataset for Classical Arabic covering the complete Quran. It contains approximately 132,736 annotated tokens organized in an extended CoNLL-X tabular structure with 43 columns. The dataset integrates orthographic representations, fine-grained morphological annotations, and a complete hybrid constituency-dependency syntactic layer. It also includes auxiliary lexical resources and annotation schemas. The resource was produced through computational processing, deep-learning-based parsing, expert-informed annotation schemes, manual curation, and validation, and is intended for research in Arabic NLP, parsing, morphology, diacritization, linguistics, digital humanities, and language technologies. | |
| dc.description.sponsorship | Princess Nourah bint Abdulrahman University Researchers Supporting Project Number PNURSP2025R263 | |
| dc.format | Text | |
| dc.format | Tabular data | |
| dc.format | Annotated linguistic corpus | |
| dc.format | Processed data | |
| dc.format | Analyzed data | |
| dc.format.extent | Approximately 132,736 tokens; extended CoNLL-X format; 43 columns; auxiliary lexicons and annotation schemas | |
| dc.identifier.citation | Nashir, W. A., Mohsen, A. M., Al-Shargabi, A. A., Nour, M. K., & Al-Onazi, B. B. (2025). Quranic [Data set]. Mendeley Data, Version 1. https://doi.org/10.17632/rk96pn66m4.1 | |
| dc.identifier.doi | https://doi.org/10.17632/rk96pn66m4.1 | |
| dc.identifier.uri | https://repository.ust.edu.ye/handle/123456789/2770 | |
| dc.language.iso | ar | |
| dc.language.iso | en | |
| dc.publisher | Mendeley Data | |
| dc.relation.isreferencedby | https://doi.org/10.1016/j.dib.2025.111940 | |
| dc.relation.uri | https://github.com/NoorBayan/Quranic | |
| dc.relation.uri | https://github.com/NoorBayan/Noor | |
| dc.rights | Creative Commons Attribution 4.0 International (CC BY 4.0). Third-party components remain subject to their original licenses and attribution requirements. | |
| dc.rights.uri | https://creativecommons.org/licenses/by/4.0/ | |
| dc.source | Mendeley Data | |
| dc.source.uri | https://data.mendeley.com/datasets/rk96pn66m4/1 | |
| dc.subject | Holy Quran | |
| dc.subject | Classical Arabic | |
| dc.subject | Quranic Treebank | |
| dc.subject | Treebank | |
| dc.subject | Natural Language Processing | |
| dc.subject | Computational Linguistics | |
| dc.subject | Morphological Analysis | |
| dc.subject | Syntactic Analysis | |
| dc.subject | Dependency Parsing | |
| dc.subject | Constituency Parsing | |
| dc.subject | Artificial Intelligence | |
| dc.subject | Machine Learning | |
| dc.subject | Corpus Linguistics | |
| dc.title | Quranic: Extended Quranic Treebank (EQTB) Dataset | |
| dc.title.alternative | Quranic | |
| dc.type | Dataset |
ملفات
الحزمة الرئيسية
1 - 5 من 12