Opportunity brief
About the Role
We are looking for an AI Engineer to build intelligent systems that automatically discover, collect, organize, and enrich AI-related information from across the web. You will develop scalable data acquisition pipelines, AI-powered enrichment workflows, and search infrastructure that powers a structured knowledge platform.
This role combines AI engineering, backend development, automation, web crawling, and data engineering. You will work on everything from acquiring data to transforming it into high-quality, searchable knowledge using Large Language Models (LLMs), vector search, and modern AI technologies.
Responsibilities
- Build AI-powered data acquisition, enrichment, and automation systems.
- Design and develop scalable web crawlers, APIs, and data pipelines.
- Build intelligent knowledge systems using LLMs, embeddings, vector databases, and AI agents.
- Develop semantic search, recommendation, and discovery capabilities.
- Own the end-to-end lifecycle of data collection, processing, validation, and indexing.
- Build scalable crawlers and automated pipelines to collect data from websites, APIs, documentation, GitHub, research portals, directories, blogs, and other public sources.
- Discover and maintain structured information on AI tools, models, datasets, research papers, benchmarks, APIs, MCP servers, companies, and emerging technologies.
- Use LLMs to extract, summarize, classify, tag, and enrich structured and unstructured data.
- Build systems for data validation, deduplication, entity resolution, and metadata enrichment.
- Design semantic search, recommendation engines, vector search, and knowledge graph pipelines.
- Develop backend services and APIs for data ingestion, processing, indexing, and retrieval.
- Monitor crawler health, automate recurring data updates, and ensure data freshness and quality.
- Integrate third-party APIs and external data sources while optimizing performance, scalability, and reliability.
- Research and evaluate new AI technologies, models, and data sources to continuously improve platform intelligence.
- Collaborate with Product and Engineering teams to build reliable, scalable, and AI-native data infrastructure.
Requirements
- Strong proficiency in Python and backend development.
- Experience with Playwright, Scrapy, Selenium, BeautifulSoup, or similar web crawling frameworks.
- Experience building scalable ETL/data ingestion pipelines and REST APIs.
- Strong understanding of HTML, JSON, APIs, web technologies, and data extraction techniques.
- Experience with PostgreSQL, MongoDB, Redis, Elasticsearch/OpenSearch, or similar databases.
- Familiarity with LLMs, AI APIs, embeddings, vector databases, Retrieval-Augmented Generation (RAG), and AI agents.
- Experience with Git, Docker, and modern cloud development practices.
- Strong analytical, debugging, and problem-solving skills.
- Self-driven with the ability to work in a fast-paced startup environment and take ownership of end-to-end engineering problems.
- Bonus: Experience with Knowledge Graphs, Hugging Face, GitHub APIs, LangChain, LlamaIndex, MCP ecosystem, Airflow, Temporal, or distributed systems.