[my talk went pretty well. whee! freedom for the rest of the conf! love when I’m up early.]
Jane Morris is at the University of Toronto. The full title of her paper is “Readers subjective perceptions of lexical cohesion and its implications for computers’ interpretations of text meaning.”
[vz: interesting. I’m not sure what I think about the phrase computers’ interpretations. slippery slope from here to computer sentience, which I certainly ain’t against, but which is a controversial topic at best.]
Meaning of text can be approached from three different (and much debated) points of view: what the readers think it means (attentional structure), what we think the author(s) thought it meant (intentional structure), and what the text itself means (linguistic structure).
Machine text and corpus analysis (computational linguistics) can only access the meaning inherent [?] in the text, but not the readers’ or writers’ perspectives. Computational linguistics has often proceeded as though the attentional and intentional structures of texts didn’t exist.
Properties of text, according to David Olson: it’s an artifact; it’s a representation of knowledge/meaning; and it needs to be interpreted by readers.
Morris’ research question: how subjective is the interpretation of lexical cohesion of text? Investigated this in a study using 26 readers and 3 texts.
Some definitions. Lexical cohesion: contribution to a text’s meaning by groups of related words running through it. Linguistics studies: meaning, coherence. Computational linguistics studies: structure of text, summarization, information retrieval, spelling correction.
Study’s objectives: summarize agreement between readers on work groups, using individual diffs as an indicator of objectivity. Readers underlined groups of words they thought were related in the same color, different groups being different colors. They then wrote out those word groups separately, and wrote a brief description of what they thought the groups meant. (Funeral, communion, chapel, deceased: processes involved when someone dies.)
Study showed about 40% individual difference in readers’ lexical interpretation of texts. This implies significant limitations in machine analysis of texts. How to computationally account for differences interpretation? Morris proposes developing reader and writer models [vz: agent-based modelling might be useful?] Simple reader model: reader-specific corpora, thesauri, and view of lexical cohesion (use reader-based thesaurus – which can be generated using off-the-shelf software – to identify word groups in text).
Future research: different texts and different readers; subjectivity of other aspects of meaning; what does subjectivity mean or reflect (reader attitudes); [more on] how to create reader models, and also writer models to reflect writer subjectivity.
