Friday, November 26, 2010

TREC 2010 Roundup

Back from another successful TREC conference on the NIST campus. 2010 is a transition year, with the end of old tracks and the proposition of new ones. Indeed, TREC is moving with the times, looking at new data sources and test collections, as well as new evaluation strategies.

Outwith the old . . .

For example, TREC 2010 marks the end of the Relevance Feedback and Blog tracks. While TREC 2010 will be the last year of the Relevance Feedback track, the Blog track, which has been running for the last 5 years, is now morphing into a new Microblog track, investigating real-time and social search tasks in Twitter. A brand new test collection possibly containing 2 months of tweets is planned, with linked web-pages and a partial follower graph. Join the Microblog track googlegroup to obtain the latest updates and follow the Microblog track on Twitter.

TREC 2011 will also witness the initiation of the new Medical Records track, dedicated to investigating approaches to access free-text fields of electronic medical records.

On the test collection front, the Web track is also forward planning a new large-scale dataset to replace ClueWeb09. Indications are that this new dataset will be about the same scale as ClueWeb09 but might provide more temporal information (multiple versions of a page or site over time). Moreover, we have suggested that this might be the heart of a larger dataset comprised of multiple parallel/aligned corpora, for example blogs and news feeds covering the same timeframe.

TREC Assessors, Relevant?

In terms of evaluation, 2010 marks the first year where evaluation judgments were crowdsourced using an online worker marketplace, as opposed to relying on TREC assessors, the participants themselves, or a select group of experts. Indeed, both the Blog track and the Relevance Feedback track crowdsourced some of their evaluation (although the Relevance Feedback track suffered many setbacks and its crowdsourcing process is still incomplete). Furthermore, to investigate the challenges in this new field of crowdsourcing, a specific Crowdsourcing track has been created and will run in 2011. More details can be found here.

Themes

As usual, themes emerged within the various tracks. Learned approaches were far more prevalent this year, now that training data was available for the ClueWeb09 dataset. Indeed, the Web track was dominated by trained models mostly based on link and proximity search features. Diversification, on the other hand, remains a challenging task, with the top groups leaving their initial rankings as is. An outstanding exception is our own approach using the xQuAD framework under a selective diversification regime, which further improves our strongly performing adhoc baseline. Craig Macdonald presented our work in the Web track plenary session.

In the Blog track, voting model-based and language modeling approaches proved popular for blog distillation. For faceted blog ranking, participants employed variants of facet dictionaries to either train a classifier or as features for learning. For the top news task, participants deployed a wide variety of methods to rank news stories in a real-time setting, from probabilistic modeling to blog post voting with historical evidence. Richard Mccreadie presented our work on the blog track as a poster during TREC 2010, which attracted very interesting discussions.

During the TREC conference, Iadh Ounis, Richard Mccreadie and others have done a fair amount of tweeting. You can follow some bits of the TREC conference through the #trec2010 hashtag.

Wednesday, November 3, 2010

CIKM 2010 in Toronto, ON, Canada

I'm back from Toronto, where a few of us attended the CIKM 2010 conference last week. On Friday, I presented our paper on "Selectively diversifying Web search results", a joint work with Craig Macdonald and Iadh Ounis. This work extends our successful participation in the diversity task of the TREC 2009 Web track, by investigating the need for search result diversification in the first place. In particular, we proposed a novel supervised learning approach to predict not only whether promoting diversity is beneficial, but also how much diversification should be applied to attain an effective retrieval performance on a per-query basis. After thorough, large-scale experiments with over 900 query features, we found that our selective approach can substantially improve existing diversification approaches, including our state-of-the-art xQuAD framework. Nonetheless, we believe the significance of our contribution goes beyond these successful results. Indeed, it was with great pleasure that we heard from the NTCIR organisers that NTCIR-9 will run an Intent task, aimed---among other things---at selectively diversifying search results, an area where we are proud to be pioneers.
Besides our own paper, a few other papers caught my attention:
  • Web Search Solved? All Result Rankings the Same? by Hugo Zaragoza, B. Barla Cambazoglu and Ricardo Baeza-Yates
  • Reverted Indexing for Feedback and Expansion, by Jeremy Pickens, Matthew Cooper and Gene Golovchinsky
  • Rank Learning for Factoid Question Answering with Linguistic and Semantic Constraints, by Matthew Bilotti, Jonathan Elsas, Jaime Carbonell and Eric Nyberg
  • Organizing Query Completions for Web Search, by Alpa Jain and Gilad Mishne
  • Clickthrough-Based Translation Models for Web Search: from Word Models to Phrase Models, by Jianfeng Gao, Xiaodong He and Jian-Yun Nie
The conference also featured great keynotes, of which those by Jamie Callan and Susan Dumais deserve a particular mention. Jamie talked about his view for the future of search, in which search engines capable of fully leveraging the structure of queries and documents would enable more sophisticated applications built on top of them. Susan addressed the temporal evolution of Web content, how it impacts the way users access this content, and how test collections should account for it. For more details, have a look at the excellent posts by Gene Golovchinsky on Jamie and Susan's talks.
Last but not least, many of us were involved in promoting the next edition of CIKM, to be held here in Glasgow. There was a lot of excitement from the several people that visited our booth, and also during the hand-over talk at the end of the conference. Well done Jon, Mary, Craig, and Iadh for the hard work! The arrangements for CIKM 2011 are well advanced, and the call for papers is now online. You can also follow the latest news about CIKM 2011 on Twitter, Facebook, LinkedIn, and Lanyrd. We look forward to welcoming you all to Glasgow next year!

Tuesday, July 20, 2010

Terrier Team at SIGIR 2010 in Geneva

SIGIR 2010 has just started in Geneva. From the TerrierTeam, Richard and myself are attending.

On Monday, Richard presented his PhD topic, Leveraging User-generated Content for News Search at the doctoral consortium.

Later, at the Web Ngram workshop, I'll be presenting a paper on Global Statistics in Proximity Weighting Models.

About the same time, Richard will be presenting at the Crowdsourcing for Search Evaluation workshop. His paper on Crowdsourcing a News Query Classification Dataset examines the effectiveness of different interfaces for having Mechanical Turkers classify queries as news-related or not.

Last but not least, and continuing on our proximity theme, Nicola Tonellotto from CNR is presenting our joint work titled Efficient Dynamic Pruning with Proximity Support at the Large Scale & Distributed Systems workshop.

Meanwhile, please say hello if you see us at the conference, or stay up to date by following #sigir2010. And remember, if you are near the registration desk, please pick up flyers for Terrier and CIKM 2011.

Monday, July 19, 2010

Top Authors in Information Retrieval

Thanks to Sérgio Nunes who alerted us to this ranking by Microsoft Academic Search of the Top Authors in Information Retrieval, in the past 5 years.

According to this recent ranking, two members of the TerrierTeam, namely Iadh Ounis and Craig Macdonald, are in the top 5 authors in Information Retrieval in the past 5 years (position #1 and #4, respectively). The ranking is based on in-domain citations.

This good news comes just at the start of the SIGIR 2010 Conference, which will be held in Geneva, Switzerland this week (19-23 July 2010). Several members of the team will be in attendance.

Tuesday, May 4, 2010

WWW 2010 in Raleigh, NC, USA

I am back from the sunny Raleigh, NC, USA. Besides the nice weather, I had a great time last week attending the 19th International World Wide Web Conference (WWW 2010), where I presented our paper on Exploiting query reformulations for Web search result diversification, a joint work with Craig Macdonald and Iadh Ounis. The paper introduces a probabilistic formulation of our xQuAD framework for search result diversification, and analyses the effectiveness of query reformulations provided by three commercial search engines for the diversification task. My talk was very well received, with lots of questions from the audience, and subsequent chatting with many people from both academia and industry.

The blend academia-industry was indeed a signature of WWW. I was also impressed with the multidisciplinary nature of the confere
nce—with up to five parallel sessions, there was always something for everyone! In particular, from the sessions I attended, a few papers caught my attention:
  • Clustering query refinements by user intent, by Eldar Sadikov et al. (Stanford University and Google)
  • Optimal rare query suggestion with implicit user feedback, by Yang Song and Li-wei He (Microsoft Research)
  • Building taxonomy of Web search intents for name entity queries, by Xiaoxin Yin and Sarthak Shah (Microsoft Research)
  • Exploring Web scale language models for search query processing, by Jian Huang et al. (Microsoft Research Asia, Facebook, and Penn State University)
  • Classification-enhanced ranking, by Paul N. Bennett et al. (Microsoft Research)
  • Ranking specialization for Web search: A divide-and-conquer approach by using topical RankSVM, by Jiang Bian et al. (Georgia Tech and Yahoo! Labs)
  • Generalized distances between rankings, by Ravi Kumar and Sergei Vassilvitskii (Yahoo! Research)
  • Relational duality: Unsupervised extraction of semantic relations between entities on the Web, by Danushka T. Bollegala et al. (University of Tokyo)
The conference also featured three passionate keynotes:
  • Vint Cerf discussed a broad range of topics of interest on today's Web, where everything is connected: 1.8 billion users, around a billion Web-enabled mobile devices, and still a large room for growth in developing countries. Touched points included the implications of the explosion of data production on mobility, accessibility, security and privacy, intellectual property, digital preservation, as well as new technologies (e.g., cloud computing).
  • dannah boyd discussed privacy implications of the availability of "big data". Her keynote revolved around common misconceptions associated with the analysis of data produced by online social activities, as well as ethical concerns related to using this data in the first place, "just because it is accessible".
  • Carl Malamud from public.resource.org described his experiences trying to convince seven bureaucratic institutions to make public data publicly accessible. His keynote was organised around "10 rules for radicals", a guide on how to break the barriers towards negotiating with bureaucrats.
On Thursday night, the conference banquet featured an exciting performance by the North Carolina string band Carolina Chocolate Drops. Check out Snowden's Jig (Genuine Negro Jig) and Don't get trouble in your mind for a taste.
Friday held the closing ceremony, with the announcement of the award winners.
Best Paper:
  • Factorizing personalized Markov chains for next-basket recommendation, by Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme (Osaka University and University of Hildesheim)
Best Student Paper:
  • Privacy wizards for social networking sites, by Lujun Fang and Kristen LeFevre (University of Michigan)
Best Posters:
  • How much is your personal recommendation worth, by Paul Dütting, Monika Henzinger and Ingmar Weber (EPFL Lausanne, University of Vienna, and Yahoo! Research)
  • SourceRank: Relevance and trust assessment for deep Web sources based on inter-source agreement, by Raju Balakrishnan and Subbarao Kambhampati (Arizona State University)
The closing ceremony also featured a short presentation of WWW 2011, to be held in Hyderabad, India. WWW 2012 will take place in Lyon, France.

Finally, on Saturday, the IW3C2 announced the Brazilian bid as the winner to host WWW 2013, which I was very glad to hear about!