Anime Graph

AboutMarch 2026

07· Data Engineering, Software EngineeringPersonal project

A service that pulls AniList's entire catalog and rebuilds it as a connected network of anime, staff, studios and characters, so a question like "which studio has this director worked at before" becomes one query instead of dozens of lookups.

ETL

Source ↗

Write-up

AniList's public catalog gives you one anime at a time, with all its related people (staff, cast, studios) bundled inside that single record as one big nested list. That makes a question like "which studio has this director worked at before" surprisingly hard to answer: nothing about how those people and studios connect to each other is stored anywhere, only implied across many records, so answering it means pulling dozens of them and cross-referencing by hand. The source models the world as documents while the interesting questions are about relationships, and that mismatch is the same one behind recommendation systems, supply-chain mapping, fraud rings and organizational analytics.

This project pulls AniList's whole catalog automatically, page by page, then rebuilds it as a network database (Neo4j) where every anime, staff member, studio and character is its own point, connected to the others by labeled relationships like "worked on", "voiced" or "produced by". Structured that way, a question that used to take dozens of separate lookups becomes a single query. It also cleans up messy staff-role labels from the source (things like "Director", "Script" or "Chief Animation Director") into a small, consistent set of relationship types, so one job spelled a dozen different ways does not fragment into a dozen different relationships.

Pulling the catalog runs in the background, so starting a sync of several thousand anime does not freeze the rest of the app, and it slows down and retries on its own whenever AniList's API asks it to rather than getting blocked. Every full sync safely clears and rebuilds the database's internal consistency rules, so a change to the data's structure never needs a manual cleanup step, and a sync can be re-run without duplicating anything. Those are the parts that make it a harvester someone could actually point at a live service.

Things to underline
  • Pulls AniList's entire catalog automatically, page by page, slowing down and retrying on its own whenever the API asks it to rather than getting blocked
  • Rebuilds the catalog as a connected network (anime, staff, studios, characters, and more, all linked) so relationships that used to take many separate lookups become a single query
  • Cleans up inconsistent staff-role labels from the source data into one consistent set of relationship types
  • Runs syncing in the background so kicking off a multi-thousand-anime pull never freezes the app, and a sync can always be safely re-run without creating duplicates
Built with
PythonFastAPINeo4jDocker