learning-to-rank · vu amsterdam · team of three
Expedia Hotel Booking Analytics
Reordering hotel search results so the one you book lands on top. 4.96M records, 146 features, top 15% of 110 teams.
Key numbers
- 4.96Mhotel-option records
- 199,795searches
- 146features from 54 columns
- 0.418private NDCG@5
- 16 / 110final rank, top 15%
what i did
What I did
Analysed 4.96M hotel-option records across 199,795 searches and surfaced booking, conversion, pricing and position-bias patterns.
Engineered 146 analytical features from 54 raw columns covering price, position and search context, in a three-person team.
Evaluated KNN and learning-to-rank ensembles with temporal validation. Ranked 16th of 110 teams (top 15%) with a private NDCG@5 of 0.418, 2.1% below the top score.
How it works.
The problem
Someone searches for a hotel and gets back a list of properties. The job is to reorder that list so the one they end up booking sits at the top, with the ones they click right behind it. Scoring is NDCG@5, so what matters is getting the booked hotel into the first few slots.
Features that compare, not just describe
The strongest features compare each hotel to the others in the same search: how its price or location score sits against the search average. A $150 hotel is cheap in one search and expensive in the next, and the model cares about that difference more than the raw number.
Taking position bias seriously
Hotels near the top of a list get clicked more whether they are good or not. About 30% of searches were shown in random order, which gives a clean read on what users actually prefer, so click and booking rates were also computed from that subset to separate quality from placement.
Validation that respects time
Search ids turned out not to follow time, so validation holds out the latest searches by timestamp instead. Target encodings are computed out of fold, grouped by search, so no row's features are built from its own label.
The lesson I keep
Fixing the data pipeline beat tuning the model. The leakage fixes in the encoding did more for the score than the ensemble experiments.