RMIT University
Browse

Topic Difficulty: Collection and Query Formulation Effects

journal contribution
posted on 2024-11-02, 19:36 authored by Shane CulpepperShane Culpepper, Guglielmo Faggioli, Nicola Ferro, Oren Kurland
Several recent studies have explored the interaction effects between topics, systems, corpora, and components when measuring retrieval effectiveness. However, all of these previous studies assume that a topic or information need is represented by a single query. In reality, users routinely reformulate queries to satisfy an information need. In recent years, there has been renewed interest in the notion of "query variations"which are essentially multiple user formulations for an information need. Like many retrieval models, some queries are highly effective while others are not. This is often an artifact of the collection being searched which might be more or less sensitive to word choice. Users rarely have perfect knowledge about the underlying collection, and so finding queries that work is often a trial-And-error process. In this work, we explore the fundamental problem of system interaction effects between collections, ranking models, and queries. To answer this important question, we formalize the analysis using ANalysis Of VAriance (ANOVA) models to measure multiple components effects across collections and topics by nesting multiple query variations within each topic. Our findings show that query formulations have a comparable effect size of the topic factor itself, which is known to be the factor with the greatest effect size in prior ANOVA studies. Both topic and formulation have a substantially larger effect size than any other factor, including the ranking algorithms and, surprisingly, even query expansion. This finding reinforces the importance of further research in understanding the role of query rewriting in IR related tasks.

Funding

Advancing Analytical Query Processing with Urban Trajectory Data

Australian Research Council

Find out more...

History

Journal

ACM Transactions on Information Systems

Volume

40

Number

19

Issue

1

Start page

1

End page

36

Total pages

36

Publisher

Association for Computing Machinery

Place published

United States

Language

English

Copyright

© 2021 Copyright held by the owner/author(s). Publication rights licensed to ACM.

Former Identifier

2006115382

Esploro creation date

2022-06-22

Usage metrics

    Scholarly Works

    Exports

    RefWorks
    BibTeX
    Ref. manager
    Endnote
    DataCite
    NLM
    DC