Student Cafe Preference Segmentation
A segmentation built on revealed preference rather than assumption: real walking distances between where students live and the cafes near them feed a graph model, and a clustering step finds the groupings instead of being told what they should be.
This builds a small network connecting 20 students, 15 boarding houses (kos) and 15 real cafes in Surabaya, using their actual map coordinates so "how far" is a real walking distance rather than a stand-in number. Each boarding-house-to-cafe connection also carries a score blended from that cafe's rating and its facilities (wifi, air conditioning, 24-hour access), so being close and being good both factor into how appealing a cafe is from a given address.
On top of that network, a graph-based machine learning technique turns each student's position in it into a set of numbers capturing who and what they are near, and a clustering algorithm groups students with similar numbers automatically. The groups are never specified in advance. Nobody decides up front that there is a "students near campus" segment and a "students downtown" segment; whatever structure is there is what comes out. Shrinking those numbers to two dimensions makes the result something you can plot and read instead of an abstract table.
The clustering surfaces patterns already sitting in the data. Point Coffee ITS in Keputih comes out as the single most-visited cafe, with 4 of the 20 students living within an average 1.75 km of it. More usefully, several boarding houses that are nowhere near each other on a map turn out to share the exact same two connected cafes, genuine neighbours by preference rather than by geography, which is exactly the pattern a manual grouping would have missed and exactly what retail site selection, delivery zoning and catchment analysis exist to find.
The sample is the only part here that does not scale. Twenty students with real coordinates and real walking distances run through the same pipeline that would run on two hundred thousand.
- Built a network of real students, boarding houses, and cafes using actual map coordinates, so distance-based results reflect real walking distance, not a placeholder number
- Used a graph-based clustering pipeline (embeddings feeding K-Means) to group 20 students by cafe preference automatically, then checked the result by plotting it in two dimensions
- Designed a weighted scoring formula (60% cafe rating, 40% facilities) so being close and being well-equipped both shape which cafe a boarding house is drawn toward
- Added a stats endpoint that checks the data meets the assignment's requirements on every sync, catching a broken data load before it reaches the graph
- 4
- Maintained