Opportunity brief
About the Role We are looking for an AI Engineer to build intelligent systems that automatically discover, collect, organize, and enrich AI-related information from across the web. You will develop scalable data acquisition pipelines, AI-powered enrichment workflows, and search infrastructure that powers a structured knowledge platform. This role combines AI engineering, backend development, automation, web crawling, and data engineering. You will work on everything from acquiring data to transforming it into high-quality, searchable knowledge using Large Language Models (LLMs), vector search, and modern AI technologies. Responsibilities Build AI-powered data acquisition, enrichment, and automation systems. Design and develop scalable web crawlers, APIs, and data pipelines. Build intelligent knowledge systems using LLMs, embeddings, vector databases, and AI agents. Develop semantic search, recommendation, and discovery capabilities. Own the end-to-end lifecycle of data collection, processing, validation, and indexing. Build scalable crawlers and automated pipelines to collect data from websites, APIs, documentation, GitHub, research portals, directories, blogs, and other public sources. Discover and maintain structured information on AI tools, models, datasets, research papers, benchmarks, APIs, MCP servers, companies, and emerging technologies. Use LLMs to extract, summarize, classify, tag, and enrich structured and unstructured data. Build systems for data validation, deduplication, entity resolution, and metadata enrichment. Design semantic search, recommendation engines, vector search, and knowledge graph pipelines. Develop backend services and APIs for data ingestion, processing, indexing, and retrieval. Monitor crawler health, automate recurring data updates, and ensure data freshness and quality. Integrate third-party APIs and external data sources while optimizing performance, scalability, and reliability. Research and evaluate new AI technologies, models, and data sources to continuously improve platform intelligence. Collaborate with Product and Engineering teams to build reliable, scalable, and AI-native data infrastructure. Requirements Strong proficiency in Python and backend development. Experience with Playwright, Scrapy, Selenium, BeautifulSoup, or similar web crawling frameworks. Experience building scalable ETL/data ingestion pipelines and REST APIs. Strong understanding of HTML, JSON, APIs, web technologies, and data extraction techniques. Experience with PostgreSQL, MongoDB, Redis, Elasticsearch/OpenSearch, or similar databases. Familiarity with LLMs, AI APIs, embeddings, vector databases, Retrieval-Augmented Generation (RAG), and AI agents. Experience with Git, Docker, and modern cloud development practices. Strong analytical, debugging, and problem-solving skills. Self-driven with the ability to work in a fast-paced startup environment and take ownership of end-to-end engineering problems. Bonus: Experience with Knowledge Graphs, Hugging Face, GitHub APIs, LangChain, LlamaIndex, MCP ecosystem, Airflow, Temporal, or distributed systems.