-
Using Large Language Models for Qualitative Analysis can Introduce Serious Bias
Authors:
Julian Ashwin,
Aditya Chhabra,
Vijayendra Rao
Abstract:
Large Language Models (LLMs) are quickly becoming ubiquitous, but the implications for social science research are not yet well understood. This paper asks whether LLMs can help us analyse large-N qualitative data from open-ended interviews, with an application to transcripts of interviews with Rohingya refugees in Cox's Bazaar, Bangladesh. We find that a great deal of caution is needed in using L…
▽ More
Large Language Models (LLMs) are quickly becoming ubiquitous, but the implications for social science research are not yet well understood. This paper asks whether LLMs can help us analyse large-N qualitative data from open-ended interviews, with an application to transcripts of interviews with Rohingya refugees in Cox's Bazaar, Bangladesh. We find that a great deal of caution is needed in using LLMs to annotate text as there is a risk of introducing biases that can lead to misleading inferences. We here mean bias in the technical sense, that the errors that LLMs make in annotating interview transcripts are not random with respect to the characteristics of the interview subjects. Training simpler supervised models on high-quality human annotations with flexible coding leads to less measurement error and bias than LLM annotations. Therefore, given that some high quality annotations are necessary in order to asses whether an LLM introduces bias, we argue that it is probably preferable to train a bespoke model on these annotations than it is to use an LLM for annotation.
△ Less
Submitted 5 October, 2023; v1 submitted 29 September, 2023;
originally announced September 2023.
-
Bayesian Topic Regression for Causal Inference
Authors:
Maximilian Ahrens,
Julian Ashwin,
Jan-Peter Calliess,
Vu Nguyen
Abstract:
Causal inference using observational text data is becoming increasingly popular in many research areas. This paper presents the Bayesian Topic Regression (BTR) model that uses both text and numerical information to model an outcome variable. It allows estimation of both discrete and continuous treatment effects. Furthermore, it allows for the inclusion of additional numerical confounding factors n…
▽ More
Causal inference using observational text data is becoming increasingly popular in many research areas. This paper presents the Bayesian Topic Regression (BTR) model that uses both text and numerical information to model an outcome variable. It allows estimation of both discrete and continuous treatment effects. Furthermore, it allows for the inclusion of additional numerical confounding factors next to text data. To this end, we combine a supervised Bayesian topic model with a Bayesian regression framework and perform supervised representation learning for the text features jointly with the regression parameter training, respecting the Frisch-Waugh-Lovell theorem. Our paper makes two main contributions. First, we provide a regression framework that allows causal inference in settings when both text and numerical confounders are of relevance. We show with synthetic and semi-synthetic datasets that our joint approach recovers ground truth with lower bias than any benchmark model, when text and numerical features are correlated. Second, experiments on two real-world datasets demonstrate that a joint and supervised learning strategy also yields superior prediction results compared to strategies that estimate regression weights for text and non-text features separately, being even competitive with more complex deep neural networks.
△ Less
Submitted 11 September, 2021;
originally announced September 2021.
-
Percolating Plastic Failure as a Mechanism for Shear Softening in Amorphous Solids
Authors:
Vijayakumar Chikkadi,
Oleg Gendelman,
Valery Ilyin,
J Ashwin,
Itamar Procaccia
Abstract:
``Shear softening" refers to the observed reduction in shear modulus when the stress on an amorphous solid is increased beyond the initial linear region. Careful numerical quasi-static simulations reveal an intimate relation between plastic failure and shear softening. The attaintment of the steady-state value of the shear modulus associated with plastic flow is identified with a percolation of th…
▽ More
``Shear softening" refers to the observed reduction in shear modulus when the stress on an amorphous solid is increased beyond the initial linear region. Careful numerical quasi-static simulations reveal an intimate relation between plastic failure and shear softening. The attaintment of the steady-state value of the shear modulus associated with plastic flow is identified with a percolation of the regions that underwent a plastic event. We present an elementary ``two-state" model that interpolates between failed and virgin regions and provides a simple and effective characterization of the shear softening.
△ Less
Submitted 15 December, 2013;
originally announced December 2013.
-
On the Effect of Micro-alloying on the Mechanical Properties of Metallic Glasses
Authors:
Oleg Gendelman,
J. Ashwin,
Pankaj Mishra,
Itamar Procaccia,
Konrad Samwer
Abstract:
"Micro-alloying", referring to the addition of small concentration of a foreign metal to a given metallic glass, was used extensively in recent years to attempt to improve the mechanical properties of the latter. The results are haphazard and nonsystematic. In this paper we provide a microscopic theory of the effect of micro-alloying, exposing the delicate consequences of this procedure and the la…
▽ More
"Micro-alloying", referring to the addition of small concentration of a foreign metal to a given metallic glass, was used extensively in recent years to attempt to improve the mechanical properties of the latter. The results are haphazard and nonsystematic. In this paper we provide a microscopic theory of the effect of micro-alloying, exposing the delicate consequences of this procedure and the large parameter space which needs to be controlled. In particular we consider two very similar models which exhibit opposite trends for the change of the shear modulus, and explain the origins of the difference as displayed in the different microscopic structure and properties.
△ Less
Submitted 19 September, 2013;
originally announced September 2013.