Data from: Public perceptions of children's reading, literacy, and cultural representation in India: Survey findings from 346 respondents in India and the Indian Diaspora
Data files
Jul 14, 2026 version files 25.47 KB
-
codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv
3.19 KB
-
codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv
3.16 KB
-
codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv
2.63 KB
-
codebooks_perceptionsurvey_(2)_-_Residence_rules.csv
1.41 KB
-
codebooks_perceptionsurvey_(2)_-_Verification.csv
635 B
-
README.md
14.44 KB
Aug 07, 2026 version files 188.72 KB
-
codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv
3.19 KB
-
codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv
3.16 KB
-
codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv
2.63 KB
-
codebooks_perceptionsurvey_(2)_-_Residence_rules.csv
1.41 KB
-
codebooks_perceptionsurvey_(2)_-_Verification.csv
635 B
-
Public_perceptions_childrens_reading_India_deidentified.csv
109.12 KB
-
Questionnaire.pdf
42.51 KB
-
README.md
26.06 KB
Abstract
Using a cross-sectional survey of 346 adults and educationists in India and across the Indian diaspora, we examine how cultural, institutional, and digital forces shape Indian children’s readings. As an institutional reality as well as a cultural space, adults connect with the childhood literacy ecosystem through shared values regarding inclusive representation, and their shared concerns over digital addiction and AI content reveal a broad societal consensus on the modern challenges of reading. The public views consuming children’s literature as a means to cultivating empathy and resisting rote, exam-focused schooling. However, the pressures of dealing with this evolving environment as a teacher creates a sharp divergence in how educators and non-educators perceive the normative responsibility of schools to nurture these empathetic connections. Discussions around childhood reading have therefore become a significant space through which Indians critique exam-centric educational structures, challenge the impacts of digital media, and have their own reflection on the institutional constraints that shape contemporary learning environments.
Uchchingre Project, June 2026
This repository contains the anonymised dataset, coding materials, and survey instrument for the study Public Perceptions of Children's Reading, Literacy, and Cultural Representation in India: Survey Findings from 346 Respondents in India and the Indian Diaspora (Mukherjee, Sarkar & Adhikari, 2026).
Applying the exact text mappings in codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv and codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv to the survey responses file (Public_perceptions_childrens_reading_India_deidentified.csv) reproduces every subgroup count and every Likert statistic reported in the paper.
Contents
| File | Description |
|---|---|
README.md |
This file. |
Public_perceptions_childrens_reading_India_deidentified.csv |
Survey responses (de-identified), UTF-8 CSV — the primary data file. 347 rows (346 consenting respondents + 1 non-consent, retained but excluded from analysis). The two open-ended free-text columns have been removed, and the recruitment-code field processed for privacy; see §1. |
codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv |
Category definitions and decision rules for occupational coding. |
codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv |
Every unique raw occupation entry, its assigned category, and respondent count. |
codebooks_perceptionsurvey_(2)_-_Residence_rules.csv |
Category definitions and decision rules for geographic coding. |
codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv |
Every unique raw residence entry, its assigned category, and respondent count. |
codebooks_perceptionsurvey_(2)_-_Verification.csv |
Category totals (computed from the mapping files) checked against the figures reported in the paper. |
Questionnaire.pdf |
The English-language survey instrument as administered. Contains the consent form, all 26 Likert items, and the open-ended questions. Hindi and Bengali translations used for in-person administration are available from the corresponding author on request. |
1. Survey responses file (de-identified)
Public_perceptions_childrens_reading_India_deidentified.csv (primary) ·
- Size: 347 data rows, 34 columns, plus one header row. The xlsx copy retains the sheet name
Form Responses 1. - Encoding / format: UTF-8 CSV (comma-delimited, single header row); an unformatted xlsx copy is also provided. Presentational formatting from the earlier version — a frozen header row, an Excel table style, and its filter drop-downs — has been removed. No merged cells, comments, hyperlinks, formulas, or highlighting were present, and none of the removed formatting is needed to interpret or re-analyse the data.
- Provenance: Direct export from the Google Form used for online administration. In-person and telephonic responses were entered into the same form by field researchers, so the dataset is unified.
- De-identification: The two open-ended free-text columns from the original export (columns 34 and 35, corresponding to the two open-ended questions in Section 6 of the questionnaire) have been removed in full from the deposited file. In the original export these carried 37 and 33 substantive responses respectively; because free-text responses can incidentally include identifying details (named organizations, institutions, places, or distinctive phrasing), the entire columns were dropped rather than paraphrased. In the remaining recruitment-source column (now column 34), a SurveyCircle recruitment link and redemption code that appeared in the original column header were removed and the header renamed
Did you find us through SurveyCircle?; the Yes/No responses are unchanged. All quantitative columns (consent, demographics, and the 26 Likert items) are byte-for-byte unchanged, so the reproduction workflow below is unaffected.
Columns
| # | Column | Type | Notes |
|---|---|---|---|
| 1 | Timestamp | Datetime | Auto-recorded submission time. |
| 2 | Consent statement | Categorical | Yes (n = 346) or No (n = 1). One respondent selected Yes, No; treated as consenting per the paper. Non-consenting row retained in the file but excluded from all analyses. |
| 3 | Age | Categorical | 18–24, 25–34, 35–44, 45–54, 55–64, 65+. One entry recorded as 33-44 is treated as 35–44 in the 18–34 vs 35+ split. |
| 4 | Gender | Categorical | Female, Male, Prefer not to say, Other. |
| 5 | Residence | Free text | State, city, or locality. Coded via codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv. |
| 6 | Education | Free text / categorical | Ranges from No education to Doctorate. Case variants (e.g. nO EDUcation, matriculation vs Matriculation) preserved as entered. |
| 7 | Occupation | Free text | Coded via codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv. |
| 8–33 | Q1–Q26 (Likert items) | Categorical | Five-point scale: Strongly disagree, Disagree, Neutral, Agree, Strongly agree. Question text is the column header. |
| 34 | Recruitment source | Categorical | Yes / No (blank for non-responses). Header renamed to Did you find us through SurveyCircle?; the SurveyCircle link and redemption code that were in the original header have been removed. Not analytical. |
The two open-ended free-text questions administered as part of Section 6 of the questionnaire (columns 34 and 35 in the original Google Form export) have been removed from this file for de-identification. The questions themselves are documented in Questionnaire.pdf and in §3 of this README, but no response-level text is included in the deposit.
Filtering to the analytical sample
To reproduce the analyses:
keep rows where consent contains "Yes" → n = 346
Likert coding
For every quantitative reproduction in the paper, the five ordinal levels are mapped to 1–5:
Strongly disagree = 1
Disagree = 2
Neutral = 3
Agree = 4
Strongly agree = 5
Matching is case-insensitive on both the level names and the Yes/No in the consent column.
2. Codebook files
Five CSV files documenting how free-text Occupation and Residence entries were categorised, plus an internal verification file. Applying the exact text mappings in the two _mapping.csv files to the responses file reproduces every subgroup count in the paper.
codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv
14 data rows — one per occupation category. Documents the coding scheme; contains no respondent-level data.
| Column | Type | Description |
|---|---|---|
| Category | Text (categorical) | Name of the occupation category. The 14 permitted values are the complete, mutually exclusive set used throughout the occupation codebook: Teacher / Educator · Farmer · Student · Healthcare Professional · Corporate / IT / Engineering · Manual Labour / Domestic Worker · Self-employed / Business · Creative Professional · NGO / Social Work · Writer / Media · Legal Professional · Homemaker · Retired · Other / Unclassifiable. |
| Definition / inclusion rule | Text (free) | The rule applied to decide whether a raw entry belongs in this category, including illustrative raw entries. |
| Key decisions and boundary cases | Text (free) | Ambiguous entries and the reasoning behind their placement. Contains -- where not applicable — see Missing values and special codes below. |
codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv
83 data rows — one per unique raw occupation string. The n respondents column sums to 346.
| Column | Type | Unit | Description |
|---|---|---|---|
| Raw occupation entry (verbatim) | Text (free) | — | The occupation string exactly as typed by respondents, preserving original capitalization, spelling, and abbreviations. Case variants of the same word appear as separate rows (e.g. Teacher and teacher), because this column reproduces the raw export rather than a cleaned version of it. The literal string (blank) is a special code — see below. |
| Assigned category | Text (categorical) | — | The occupation category assigned to this entry. Takes one of the 14 values listed in codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv. |
| n respondents | Integer | Respondents | Number of respondents who gave this exact raw entry. Range 1–93. Sums to 346 across all rows. |
codebooks_perceptionsurvey_(2)_-_Residence_rules.csv
4 data rows — one per residence category. Documents the coding scheme; contains no respondent-level data.
| Column | Type | Description |
|---|---|---|
| Category | Text (categorical) | Name of the residence category. The 4 permitted values are the complete, mutually exclusive set used throughout the residence codebook: West Bengal · Other Indian state/UT · India (state not specified) · Outside India. |
| Definition / inclusion rule | Text (free) | The rule applied to decide whether a raw entry belongs in this category, including illustrative raw entries. |
| Key decisions and boundary cases | Text (free) | Ambiguous entries and the reasoning behind their placement. Complete for all 4 rows; no missing values. |
codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv
81 data rows — one per unique raw residence string. The n respondents column sums to 346.
| Column | Type | Unit | Description |
|---|---|---|---|
| Raw residence entry (verbatim) | Text (free) | — | The residence string exactly as typed by respondents, preserving original capitalization, spelling variants, and misspellings (e.g. hoogly, Andra Pradesh). Entries range from state names to city names to individual Kolkata localities. Case variants appear as separate rows. |
| Assigned category | Text (categorical) | — | The residence category assigned to this entry. Takes one of the 4 values listed in codebooks_perceptionsurvey_(2)_-_Residence_rules.csv. |
| n respondents | Integer | Respondents | Number of respondents who gave this exact raw entry. Range 1–103. Sums to 346 across all rows. |
codebooks_perceptionsurvey_(2)_-_Verification.csv
20 rows: 14 occupation categories, 1 occupation total, 4 residence categories, 1 residence total. This is an internal consistency check, not source data.
| Column | Type | Unit | Description |
|---|---|---|---|
| Category | Text (categorical) | — | An occupation or residence category name, or one of two summary rows. Rows 1–14 are occupation categories; row 15 is Total occupation; rows 16–19 are residence categories; row 20 is Total residence. The two Total rows are sums, not categories, and must be excluded when treating this file as a category-level table. |
| Count (formula from mapping sheets) | Integer | Respondents | Respondents in that category, obtained by summing n respondents over all rows of the corresponding mapping file that carry that category. In the source workbook this was a live formula; in the CSV export it is the computed value. |
| Figure reported in paper | Integer | Respondents | The corresponding figure as printed in the manuscript, transcribed manually. |
| Match | Text (categorical) | — | Result of comparing the two preceding columns. OK = the values are equal. All 20 rows read OK. |
Scope of this check. This file verifies that the figures printed in the paper match the totals computed from the mapping files. It cannot detect an error in the mapping itself, since an entry assigned to the wrong category will still be counted consistently in both places. The _rules.csv files, not this one, are the record of coding decisions.
General coding principles
- Coding is based only on the exact text entered; no cross-field inference.
- Matching ignores capitalisation and leading/trailing whitespace (the raw export contains answers with varied capitalisation and spacing).
- Each respondent is assigned exactly one category per codebook.
- Entries naming a former occupation (e.g.
Retired teacher) are coded by the named occupation; bareRetiredwith no occupation named is coded Retired. - Entries too vague or unusual to classify are coded Other / Unclassifiable (occupation) or by the most specific identifiable location (residence).
Missing values and special codes
There are no empty cells, and no n/a or NA infills, in any of the five CSV files. Two special codes are used, both intentionally:
-- in codebooks_perceptionsurvey_(2)_-_Occupation_rules.csv, column Key decisions and boundary cases
Meaning: not applicable. No boundary cases or contested entries arose for that category, so there was nothing to record. It does not mean a value is unavailable, missing, withheld, or pending; there is no absent information behind these cells.
-- appears on 5 of the 14 rows: Self-employed / Business, NGO / Social Work, Legal Professional, Homemaker, and Other / Unclassifiable. For the first four, every raw entry fell unambiguously inside the category definition. For Other / Unclassifiable, the boundary cases are documented in the rows of the categories they were considered for and excluded from, rather than duplicated here.
The equivalent column in codebooks_perceptionsurvey_(2)_-_Residence_rules.csv is complete for all 4 rows and contains no --.
(blank) in codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csv, column Raw occupation entry (verbatim)
Meaning: the respondent left the occupation field empty. One respondent submitted no occupation. Because this file is keyed on the verbatim raw string, that empty response is represented by the literal placeholder text (blank) rather than an empty cell, so that the row remains visible and countable. It is coded Other / Unclassifiable and carries n respondents = 1. This is the only respondent with no occupation recorded, and the only such placeholder in the deposit.
No comparable placeholder exists in codebooks_perceptionsurvey_(2)_-_Residence_mapping.csv: all 346 respondents gave a residence.
3. Questionnaire
Questionnaire.pdf — the English-language survey instrument as administered.
Contains six sections, in order:
- Consent — statement and eligibility check (age ≥ 18); respondents choose Yes, I Agree or No, I do not Agree.
- Demographic Information — age, gender, residence, highest level of education completed, occupation. All five fields are free-response in the PDF; on the online form they were rendered as a mix of dropdowns and short-answer fields.
- Reading Habits, Literacy Perceptions & Diversity — 13 Likert items covering literacy, publishing, and cultural representation.
- AI Content and Digital Media Consumption — 5 Likert items.
- Empathy and socio-cultural understanding in children — 8 Likert items.
- Open-ended Questions & Survey Wrap-up — two open-ended questions plus one administrative question on whether the respondent came via Survey Circle.
All 26 Likert items use the same five-point scale: Strongly Disagree · Disagree · Neutral · Agree · Strongly Agree.
Mapping to the paper's analytical structure
The paper analyses the 26 Likert items in five thematic subscales, which regroup Section 3 of the PDF into three narrower blocks:
| Paper subscale | Items | Located in PDF |
|---|---|---|
| Literacy and Reading Habits | Q1–Q4 | Section 3 |
| Children's Publishing Landscape | Q5–Q8 | Section 3 |
| Cultural Representation in Literature | Q9–Q13 | Section 3 |
| Digital Media and AI-Generated Content | Q14–Q18 | Section 4 |
| Empathy and Reading | Q19–Q26 | Section 5 |
Item order in the PDF matches item order in the raw responses file (columns 8–33) and the Q-numbering used throughout the paper.
Open-ended questions
Two of the three questions in Section 6 are analytical:
- What do you think is the biggest challenge facing children's reading culture in India today? What can we do to mitigate those challenges?
- How do we ensure cultural diversity and representation of marginalized communities in children's literature?
Language versions
The PDF is the English-language version. The instrument was translated into Hindi and Bengali for in-person administration to daily-wage labourers, farmers, domestic helpers, and homemakers in West Bengal and in urban informal settlements. The translated versions are available from the corresponding author on request.
Reproducing the paper's numbers
A minimal reproduction workflow:
- Load
Public_perceptions_childrens_reading_India_deidentified.csv - Keep rows where the consent column contains
Yes→ n = 346. - For occupation and residence subgroup counts, join on the exact raw text against
codebooks_perceptionsurvey_(2)_-_Occupation_mapping.csvandcodebooks_perceptionsurvey_(2)_-_Residence_mapping.csv, then aggregate byAssigned category. - For every Likert item Q1–Q26, map levels to 1–5 as specified above, then compute frequency counts, combined agree% (SA + A), mean, and SD. These reproduce Tables 2–6.
- For inter-item correlations, use Spearman's rho on the 1–5-coded Likert items. These reproduce the correlations reported in Section 3.6 and the Appendix.
- For teacher / non-teacher and farmer / non-farmer comparisons, use Mann–Whitney U. For four-group occupation comparisons, use Kruskal–Wallis H.
Cross-checking against codebooks_perceptionsurvey_(2)_-_Verification.csv to confirm that the subgroup base counts (teachers n = 107, farmers n = 63, West Bengal n = 219, etc.) have been reproduced correctly before running any tests.
Ethics and consent
The study was conducted by independent researchers unaffiliated with an institution, so no institutional ethics committee review was available. All participants were 18 years or older, provided informed consent by actively selecting Yes on the consent statement before beginning the survey, and were free to withdraw at any time. No personal identifiers were recorded beyond general demographic categories (age range, gender, approximate residence, education level, employment status), and all data are reported in aggregate. Prior to deposit, the two open-ended free-text response columns were removed in full to eliminate any residual re-identification risk from incidental identifying details (such as named organizations or places) in participants' own wording, and a recruitment-platform link and code were removed from the recruitment-source column header.
Citation
Mukherjee, S., Sarkar, S., & Adhikari, J. (2026). Public Perceptions of Children's Reading, Literacy, and Cultural Representation in India: Survey Findings from 346 Respondents in India and the Indian Diaspora. Uchchingre Project.
Contact
Corresponding author: Suhasini Mukherjee — spockox@gmail.com.
For questions about the coding decisions in the codebook files, please contact the corresponding author directly.
Human subjects data
All participants in this study were 18 years of age or older and provided explicit, informed consent prior to participation. Participants were required to actively select "yes" on an informed consent document at the beginning of the survey, which detailed the study's purpose, their right to withdraw at any time, and their consent for the resulting de-identified data to be shared in the public domain. As this research was conducted by independent researchers, institutional ethics review board (IRB) approval was not required.
To ensure complete de-identification and protect participant privacy, no direct personal identifiers (such as names, contact information, or IP addresses) were collected during the survey. The dataset contains only generalized demographic information, specifically age range, gender, approximate residence, education level, and employment status. Where participants provided free-text answers for their occupation and geographic location, these entries were mapped and recoded into broad, aggregated categories prior to publication to prevent any possibility of indirect re-identification. No foreseeable risks are associated with this survey, and the resulting public dataset is fully anonymized.
Changes after Jul 14, 2026: Survey response and questionnaire uploaded for transparency, and in CSV format
