Voltar às notícias de pesquisa ENA
Photorealistic editorial photograph of qualitative researchers reviewing educator narratives beside transparent NLP clusters and an audit notebook
Artigo de conferênciaArtigo de anais revisado por pares202128 de jul. de 20263 min de leitura

Human–NLP collaboration made three dimensions of qualitative trustworthiness inspectable

Ha Nguyen, June Ahn, Ashlee Belgrave, Jiwon Lee, Lora Cawelti, Ha Eun Kim, Yenda Prado, Rossella Santagata, Adriana Villavicencio

2nd International Conference on Quantitative Ethnography (ICQE 2020)

O resumo revisado do artigo está disponível atualmente em inglês.

Resumo revisado

Photorealistic editorial photograph of qualitative researchers reviewing educator narratives beside transparent NLP clusters and an audit notebook

Establishing Trustworthiness Through Algorithmic Approaches to Qualitative Research, a 2021 conference paper by Ha Nguyen, June Ahn, Ashlee Belgrave, Jiwon Lee, Lora Cawelti, Ha Eun Kim, Yenda Prado, Rossella Santagata, Adriana Villavicencio, examines how natural language processing combined with researcher interpretation can strengthen credibility, dependability, and confirmability in qualitative analysis. The workflow is illustrated through a time-sensitive study of educators' experiences during the 2020 COVID-19 pandemic. The review therefore starts from the paper's actual evidence source and purpose rather than from the visual appeal of its final network.

The analytic move is important because ENA represents relations among coded elements, not merely how often each element appears. Algorithmic pattern finding and human review are documented as complementary stages, allowing researchers to inspect candidate groupings, challenge interpretations, and record how analytic decisions evolved. In a defensible workflow, units define whose or what network is accumulated, conversation boundaries and windows define where proximity can become a connection, and the coding scheme defines which aspects of the source material enter the model. Those decisions determine the estimand before normalization, projection, rotation, or plotting begins. A network line is consequently a modeled connection under a documented specification; it is not a direct photograph of thought, collaboration, identity, or learning.

The authors report that the case demonstrates practical affordances for efficient analysis while keeping human examination central to the claims and to the trustworthiness audit. This result is most useful as a relational account: it identifies which coded elements were organized together under the study's data and model choices. It should be read alongside unit-level variation, source excerpts or events, and any reported comparison statistics. Visual distance, line thickness, or an attractive subtraction network alone cannot establish practical importance. When a paper combines network output with qualitative return, experimental contrast, trace evidence, or another analytic view, those components strengthen interpretation because they make competing explanations easier to inspect.

The claim boundary is equally central. A successful illustration does not make NLP output inherently credible, transferable, or unbiased; trustworthiness still depends on data quality, researcher reflexivity, validation, and transparent disagreement handling. ENA cannot on its own repair a weak sample, an unstable codebook, missing contextual evidence, inappropriate dependence assumptions, or a window that crosses contexts that should remain separate. Nor does dimensional reduction preserve every feature of a high-dimensional connection space. The safest conclusion separates three layers: what was observed or collected, what the specified model represents, and what broader explanation the research design can support. Any transfer to a new population, language, activity, platform, or analytic pipeline requires fresh validation rather than visual analogy.

For ENA.HK readers, the paper's durable contribution is that it reframes computational assistance as part of a visible qualitative process rather than an opaque substitute for interpretation. A reproducible application should save the source-data provenance, segmentation and ordering rules, unit and conversation fields, code definitions, window and weighting choices, normalization and rotation settings, software version, exclusions, and sensitivity checks. It should also retain a route back from every interpreted edge to the qualitative excerpt, observed event, trace record, image element, or document that generated it. That evidence chain keeps the quantitative model and ethnographic meaning in deliberate contact while preventing a descriptive network pattern from being overstated as a causal or universal finding.