Wednesday, April 28, 2010

RIAO 2010 in Paris, France.

The 9th International RIAO Conference has started in Paris, France (28-30 April, 2010). It is unfortunate that it is being held concurrently with WWW 2010 in Raleigh.

The first RIAO conference was held in Grenoble in 1985. RIAO is currently a triennial conference, addressing Information Retrieval research topics of interest to both Academia and Industry. This year, the conference focuses on Adaptivity, Personalization and Fusion of Heterogeneous Information.

The following papers have caught my eyes, while browsing the RIAO 2010 program:
  • Boiling down information retrieval test collections. T. Sakai et al. (Microsoft Research Asia, CMU)
  • Improving tag recommendation using social networks. A. Rae et al. (The Open University, Yahoo! Research Barcelona).
  • Analysis of robustness in trust-based recommender systems. Z. Cheng and N. Hurley (UCD)
  • Opinion-finding in blogs: A passage-based language modelling approach. M. Saad Missen et al (IRIT)
  • Predicting query performance using query, result, and user interaction features. Q. Guo et al. (Emory University/Microsoft Research)
  • Towards a collection-based results diversification. J.A. Akinyemi et al. (University of Waterloo)
In addition, the TerrierTeam has two full papers, which are being presented today at the conference (hopefully, the slides will follow shortly):
  • Voting for Related Entities by R.L.T. Santos, C. Macdonald and I. Ounis. The paper addresses the problem of entity search, where the goal is to rank not documents, but entities in response to a given query. The paper proposes to tackle this problem as a voting process, by considering the occurrence of an entity among the top ranked documents for a given query as a vote for the existence of a relationship between this and the entity in the query. The approach led to high precision and unparalleled recall compared to TREC 2009 systems.
  • News Article Ranking: Leveraging the Wisdom of Bloggers by R. McCreadie, C.Macdonald and I. Ounis. The paper investigates how news article ranking can be performed automatically, so as to assist editors in selecting the articles, which should make the front page of their newspaper. In particular, the paper investigates the blogosphere as a prime source of evidence, on the intuition that bloggers, and by extension their blog posts, can indicate interest in one news article or another. The paper proposes to model the automatic news article ranking task as a voting process, where each relevant blog post acts as a vote for one or more news articles. The approach led to the best TREC 2009 retrieval performance in the Blog track.
Craig Macdonald is tweeting the conference, pending an appropriate wireless signal. You can follow some bits of the RIAO conference through the #riao2010 hashtag.

Wednesday, April 7, 2010

ECIR 2010 in Milton Keynes: A Report

Last week, five of us attended the ECIR 2010 conference in Milton Keynes. The conference was fairly well-organised, although it markedly lacked the lustre of the previous three editions of the conference. In terms of attendance, only about 170 delegates have registered, much less than Glasgow 2008 (210+), and Toulouse 2009 (180+). Perhaps, the exotic town of Milton Keynes was not deemed to be a very attractive venue for a conference. In fact, apart from attending the conference, there was not much else to do -- e.g. the nearest proper pub was at about 2 miles from the conference venue.

The ECIR 2010 conference has suffered from a new and previously unseen problem: several authors and presenters did not make it to the conference, preferring to give their presentation by proxy or using a pre-recorded talk. No less than 5 no-shows were recorded during the conference. Even the keynote speaker and winner of the first BCS IRSG Karen Sparck Jones award, Mirella Lapata, did not show up and gave her presentation through a pre-recorded video. While Lapata certainly had a valid reason (as probably did the other speakers) not to show up, it is clear that ECIR should concretely deal with such a problem, e.g., by making it compulsory that at least one author of each accepted paper be present during the conference.

In addition, the organisers decided not to have parallel sessions (because of lack of facilities?) during ECIR 2010. Therefore, several full papers were turned into poster presentations, which were held during the short lunch period. This was a very bad move, as because of the setting, these papers received much less attention and credit, even compared to the actual posters, the session of which was rather successful. Some delegates argued that some of the full-papers-turned-posters should have been given a full presentation slot, in lieu of those full papers with a no-show author.

Other than the problems mentioned above, the conference program was generally of a very good quality. In the first day, we enjoyed an excellent tutorial by two MSR researchers on Machine Learning for IR. The tutorial was given by Paul Bennett and Kevyn Collins-Thompson. We also enjoyed an equally excellent tutorial on Crowdsourcing by Omar Alonso from Bing.

In the next days, there were also several good papers that are worth reading:
  • A language modeling approach for temporal information needs (from Max-Planck)
  • The role of query sessions in extracting instance attributes from web search queries (from Google)
  • Interpreting user inactivity on search results (from Univ. of Washington, Univ. of Patras)
  • Learning to distribute queries onto Web search nodes (from Yahoo!)
  • Temporal shingling for version identification in Web archives (from Max-Planck)
  • Evaluation and user preference study on spatial diversity (University of Sheffield)
The best paper award was jointly awarded to:
  • Promoting ranking diversity for biomedical information retrieval using Wikipedia. Jimmy Huang and Xiaoshi Yin (York University)
  • Evaluation of an adaptive search suggestion system. Sascha Kriewel and Norbert Fuhr (University of Duisburg-Essen, Germany)
We have also had the chance to present our two full-papers on search result diversification, and learning to select:
  • Explicit search result diversification through sub-queries by Rodrygo L. T. Santos, Jie Peng, Craig Macdonald, and Iadh Ounis. Rodrygo presented our xQuAD search results diversification framework, and the talk was very well received by the delegates, leading to several questions, and many comments that this was arguably the best presentation of the conference.
  • Learning to select a ranking function by Jie Peng, Craig Macdonald and Iadh Ounis. This was one of the full-paper-turned-poster presentations. Jie presented the poster, which attracted a lot of attention and led to some very interesting discussions.

Finally, during the posters/demos session, two good contributions particularly caught our attention:
  • An Empirical Study of Query Specificity (Poster) - Avi Arampatzis and Jaap Kamps
  • NEAT :News Exploration Along Time (Demo) - Omar Alonso, Klaus Berberich, Srikanta Bedathur and Gerhard Weikum
The conference had also an Industry day, which we missed. You can see a report on the Industry day in the following blog post. During the conference, a few of us actively twittered the conference sessions. You can look at the archived ecir2010 hashtag for more details.

One of the most exciting moments of the conference was our visit to the Bletchley Park as part of the ECIR 2010 social dinner. This was an excellent venue with a lot of history, and the food was also good! During the dinner, we were given an impossible quiz to answer. Despite the wine, and a long day, some delegates did manage to find the answers.

Usually, when ECIR is held in the UK, the last day of the conference is the venue for Annual General Meeting of the BCS IRSG - the umbrella group for ECIR. However, in 2010, there was no AGM. We can only suppose that this was because the 2009 AGM was only held in October, co-located with Search Solutions 2009 at BCS HQ. We say suppose, because at the time of writing, the 2009 AGM minutes are not yet available!

Finally, we would like to thank the organisers for their hard work during the conference, for the idea of the ball-bouncer game during the session breaks, which was really cool/fun and for an overall reasonably organised conference. We look forward to ECIR 2011 in Dublin!

Wednesday, March 10, 2010

Terrier 3.0 released

Firstly, we have a new website for Terrier: http://terrier.org

Also, we have just released Terrier 3.0!

This is a major update to Terrier, including:
  • support for indexing WARC collections (such as ClueWeb09)
  • improved MapReduce mode indexing
  • improved and more scalable index structures
  • added field-based and proximity term dependence models, such as BM25F, PL2F and Markov Random Fields
  • new Web-based retrieval interface
Fuller changelog at http://terrier.org/docs/current/whats_new.html

If your looking for our team publications, etc., please see our new team website: http://terrierteam.dcs.gla.ac.uk/

Thanks are due to everyone in the Terrier Team for their hard work to make this release, as well as the contributions and feedback about Terrier from our users and collaborators.

Tuesday, February 23, 2010

TREC Blog Track 2010

The TREC Blog track will be continuing in 2010. In
 2009, 
the
 Blog 
track 
has
 been
 markedly 
revamped
, addressing 
more
 refined
 Blog 
search 
scenarios
 using 
the new Blogs08 collection, a
 large
 sample 
of
 the 
blogosphere covering the period of 14th January 2008 to 10th February 2009.

A summary of the TREC Blog track 2009 edition has been presented by Iadh Ounis at the main TREC conference (Slides). The Blog track 2009 overview paper will be available on the TREC website shortly, once it is updated and reviewed.

The details of the TREC 2010 Blog track are still being finalised by the organisers. However, following the discussions at the TREC 2009 Blog track workshop, here are some salient details (see also the TREC 2009 Wrap-up Slides):

1. Faceted blog search task will run again in 2010: The task addresses
 the 
quality aspect
 of
 the
 retrieved blogs
. It is a feed search task.
  • We will adopt a two-stage submission procedure: (1) a participating group submits "topically-relevant"blogs for each query; (2) a few standard baselines will be distributed to participants, so that they can re-rank them with respect to various facet inclinations (e.g. opinionated, in-depth, personal).
  • Groups can participate in stage 2 without stage 1, and vice-versa. Stage 1 is akin to an adhoc blog search task.
  • More topics for various facet inclinations.

2. Top news story identification task will run again in 2010: The task addresses the 
news‐related 
dimension
 of 
the 
blogosphere. In particular, it investigates whether the blogosphere can be used to identify the most important news stories of the day.


  • Real-time news search task rather than retrospective.
  • Much larger and a more comprehensive headlines sample, provided by a major news organisation.
  • A two-stage submission procedure: (1) Groups submit a ranking of top stories for some days per-category (e.g. sport, politics, business, etc.) (2) We will then select some top relevant stories, for which we will ask the participating groups to identify the related blog posts, in a manner that covers the various/diverse aspects of each story.
  • Groups can participate in stage 2 without stage 1. In the latter case, its is an adhoc diversity blog post search task, where the headline is the query.
We welcome any feedback and comments on the tasks above to trecblog-organisers (at) dcs.gla.ac.uk

Finally, note that if you wish to participate in TREC 2010, you should answer the TREC 2010 call for participation. We will update the Blog track wiki as things become more refined - keep following the Blog track developments as they happen on our dedicated Wiki web site.

Tuesday, August 4, 2009

AcademTech: Faceted People Search

AcademTech is a Computing Science-specific expert search engine based on the Terrier IR Platform. Persons working at Computing Science departments in Scottish Universities are considered as candidate experts by the system. Profiles of their expertise evidence are then mined from their homepages, publicly available digital libraries (e.g. DBLP) and related information found on the Web through Yahoo! BOSS. The ranking of experts is provided by a variant of the Voting Model expert search approach.

The system is integrated with a novel faceted search interface to allow users to browse and explore the results using a number of categories such as Location or Conference/Journal publications. Each expert in the system has a profile page containing a number of elements including query specific supporting publications, most informative associated terms displayed as a tag cloud, co-authors and web links. Although the system is currently applied in the context of Scottish Computing Science Academia, it can easily be expanded to go beyond its current Scottish scope, cover other academic fields, and people in general.

I was lucky enough to be able to demo AcademTech at SIGIR 2009 in Boston on July 20th. Thankfully, I spoke to a large number of attendees receiving largely very helpful feedback.

A popular suggestion was to utilize AcademTech's core system in the scope of biology. This would meet the medical field's need for finding related organisms, diseases etc. Possible facets in the area would likely be biological classifications such as species and genus.

Daniel Tunkelang from The Noisy Channel suggested providing profile page-located facets, allowing filtering of search results by features present in a selected expert's page such as co-authors. This would satisfy an example scenario such as "Show me co-authors of this expert who work for the University of Glasgow." Profile facets could also allow the experts publications list to be filtered by a number of fields such as co-author location, conference etc.

Much of the feedback mirrored that of intended future work. Name disambiguation is a high priority update as a current problem with AcademTech is the publication mismatch when multiple experts have the same name. In fact, the system is specifically designed to allow for expansion of facets, and name disambiguation. With a large amount of publication collaborators working in industry a useful move would be to expand to accommodate these experts.

AcademTech Sigir 2009 PosterAcademTech is now publicly accessible from http://www.terrier.org/academtech
A description of the system is available in the SIGIR'09 proceedings.

Thank you to all those who spoke to me and gave me some great feedback.