Member of Technical Staff (Search Crawler Analyst)
The internet is vast, containing trillions of URLs. Perplexity’s crawling and storage system is complex and has multiple stages (URL discovery, crawling, parsing, indexing). Each stage offers many opportunities for improvement and room for intricate bugs. You’ll work at the intersection of data analysis and engineering — designing metrics, building data pipelines, and improving the quality of our search and answer systems. Responsibilities
Find and diagnose quality issues in our crawling pipeline
Train small models that optimize particular aspects of the pipeline (e.g. parsing quality)
Build datasets for model training, including LLM-as-a-judge labeling pipelines
Improve page selection algorithms for indexing
Design and analyze experiments to validate improvements
Qualifications
4+ years of experience as a data analyst, ML engineer, or in a related role
Strong coding skills: you should be able to write production-grade code at the level of a mid-level backend engineer
Experience designing metrics from scratch
Experience training ML models that shipped to production with measurable metric improvements
Nice to have
Direct experience working on web crawling or indexing pipelines
By clicking “Accept All Cookies”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.