QURANIC / EXTENDED QURANIC TREEBANK (EQTB) README ============================================================ Dataset title: Quranic Descriptive title: Extended Quranic Treebank (EQTB): A Complete Multi-Layered Quranic Treebank Dataset with Hybrid Syntactic Annotations for Classical Arabic Version: Version 1 Original publication date: 25 April 2025 Original dataset repository: Mendeley Data Original dataset DOI: https://doi.org/10.17632/rk96pn66m4.1 Original dataset page: https://data.mendeley.com/datasets/rk96pn66m4/1 Related data article: Nashir, W. A., Mohsen, A. M., Al-Shargabi, A. A., Nour, M. K., & Al-Onazi, B. B. (2025). A complete, multi-layered quranic treebank dataset with hybrid syntactic annotations for classical arabic processing. Data in Brief, 62, 111940. https://doi.org/10.1016/j.dib.2025.111940 Institutional repository copy: University of Science and Technology Institutional Repository, Yemen. This deposit is an institutional repository copy of Mendeley Data Version 1. The original Mendeley Data DOI remains the canonical dataset DOI. ------------------------------------------------------------ 1. OVERVIEW ------------------------------------------------------------ The Extended Quranic Treebank (EQTB) is a computationally accessible, multi-layered linguistic resource for Classical Arabic covering the complete Quran. The dataset contains approximately 132,736 annotated tokens and is organized primarily in an extended CoNLL-X-style tabular structure with 43 columns. The resource integrates three principal annotation layers: 1. Orthographic layer - Uthmani and Imlaai representations - transliteration and phonetic information - English translation information - Quranic and sentence-based positional indexing 2. Morphological layer - fine-grained Part-of-Speech annotation - lemma and root information - morphosyntactic features such as case, mood, aspect, voice, person, gender, number, nominal state, prefixes and suffixes 3. Syntactic layer - hybrid constituency-dependency annotation - dependency relation labels - constituent tags and spans The dataset was developed through computational processing, deep-learning- based parsing, expert-informed annotation schemes, manual curation, and validation against authoritative Classical Arabic grammatical references. ------------------------------------------------------------ 2. FILES INCLUDED IN THIS DEPOSIT ------------------------------------------------------------ A. Quranic Primary dataset file. Contents: - approximately 132,736 annotated Quranic tokens - extended CoNLL-X-style tabular structure - approximately 43 data fields/columns - orthographic, morphological, and syntactic annotations NOTE: Preserve the original filename, extension, encoding, and internal structure of the file downloaded from Mendeley Data Version 1. B. CAMorphFeatures.xlsx Classical Arabic morphological feature schemas. Workbook sheets: - NominalCase - DerivedNoun - NominalState - Number - Prefixs - Suffixs - VerbForm - VerbMood - VerbState - VerbVoice - SpecialGroup - Gender - Person C. CAPoS.csv Part-of-Speech tag schema. Records: 45 Fields: pid, pos, pos_ar, pos_en D. RelLabels.csv Dependency/syntactic relation labels. Records: 140 Fields: rel_id, rel_en, rel_ar E. ConstituentsTags.csv Constituency tag schema. Records: 6 Fields: pid, tag, tag_ar, tag_en F. CALemmaLexicon.csv Classical Arabic lemma lexicon. Records: 4,833 Fields: lemma, lemma_ar, size G. CARootLexicon.csv Classical Arabic root lexicon. Records: 1,643 Fields: root, root_ar, size H. README.txt This documentation file. I. LICENSE.txt License and reuse notice. J. CITATION.txt Recommended citations for the dataset and associated article. K. THIRD_PARTY_NOTICES.txt Attribution and licensing notes for source resources incorporated or used during dataset development. L. FILE_MANIFEST.txt File-level inventory and technical notes. ------------------------------------------------------------ 3. IMPORTANT TECHNICAL NOTE ON AUXILIARY CSV FILES ------------------------------------------------------------ The following files have a .csv extension but are encoded as UTF-16 and use TAB characters as field delimiters: - CAPoS.csv - RelLabels.csv - ConstituentsTags.csv - CALemmaLexicon.csv - CARootLexicon.csv When importing these files into statistical, spreadsheet, or programming software, use: Encoding: UTF-16 Delimiter: TAB Do not automatically convert or overwrite the original files when preserving the deposited Version 1. CAMorphFeatures.xlsx is an Excel workbook containing multiple sheets. ------------------------------------------------------------ 4. RELATED SOFTWARE ------------------------------------------------------------ Quranic dataset/software repository: https://github.com/NoorBayan/Quranic Noor visualization/analysis software: https://github.com/NoorBayan/Noor These links are related software resources and are not substitutes for the archived Mendeley Data Version 1 files. ------------------------------------------------------------ 5. PROVENANCE ------------------------------------------------------------ Foundational Quranic text and annotation resources used during development included the Tanzil Project and the Quranic Arabic Corpus. Additional linguistic resources/texts used in the research workflow were obtained from the Comprehensive Islamic Library (Shamela). Custom Python scripts were used for processing, tokenization, transliteration, morphological re-annotation, and syntactic data preparation. A BiLSTM-based deep-learning parser using custom Word2Vec embeddings was used to generate the comprehensive syntactic layer. The resulting annotations underwent manual validation and expert review. See the associated Data in Brief article for the full methodology: https://doi.org/10.1016/j.dib.2025.111940 ------------------------------------------------------------ 6. LICENSE ------------------------------------------------------------ The Mendeley Data record for Version 1 identifies the dataset license as: Creative Commons Attribution 4.0 International (CC BY 4.0) License information: https://creativecommons.org/licenses/by/4.0/ Users must provide appropriate attribution, identify the source, and indicate changes when applicable. IMPORTANT: Some source components and foundational materials are subject to their own licenses and attribution requirements. See THIRD_PARTY_NOTICES.txt. ------------------------------------------------------------ 7. CITATION ------------------------------------------------------------ Please cite the original dataset DOI when using the data: Quranic. Mendeley Data, Version 1. https://doi.org/10.17632/rk96pn66m4.1 For a fuller contributor citation and the associated scholarly article, see CITATION.txt. ------------------------------------------------------------ 8. FUNDING ------------------------------------------------------------ Princess Nourah bint Abdulrahman University Researchers Supporting Project, Grant/Project Number: PNURSP2025R263. ------------------------------------------------------------ 9. REUSE AND REPRODUCIBILITY ------------------------------------------------------------ Potential uses include: - Classical Arabic natural language processing - dependency and constituency parsing - morphological analysis - Part-of-Speech tagging - diacritization - corpus linguistics - computational linguistics - digital humanities - Arabic language technologies - educational and pedagogical applications For reproducibility details, methods, and validation procedures, consult the associated article and the original Mendeley Data record. ------------------------------------------------------------ 10. REPOSITORY-COPY NOTICE ------------------------------------------------------------ This copy is preserved by the University of Science and Technology Institutional Repository for institutional preservation, discovery, and research reuse. It should be identified as a repository copy of the original Mendeley Data Version 1 record. The University repository Handle assigned to this copy does not replace the original dataset DOI: https://doi.org/10.17632/rk96pn66m4.1