Skip to content
← All opportunities

AI Engineer Internship

InternshipRemoteOnline
Applications closedVerify on The AI Signal →

The source listing appears to be closed. Open the company's listing to confirm — the source may still be accepting entries.

Sign in to save
Applications closed: 06/09/2026

Skills for this role

Skills we detected for this role.

  • Python
  • HTML
  • REST APIs
  • Git
  • Docker
  • PostgreSQL
  • MongoDB
  • Redis
  • Elasticsearch
  • Data Collection
  • Apache Airflow
  • Playwright

Opportunity brief

About the Role

We are looking for an AI Engineer to build intelligent systems that automatically discover, collect, organize, and enrich AI-related information from across the web. You will develop scalable data acquisition pipelines, AI-powered enrichment workflows, and search infrastructure that powers a structured knowledge platform.

This role combines AI engineering, backend development, automation, web crawling, and data engineering. You will work on everything from acquiring data to transforming it into high-quality, searchable knowledge using Large Language Models (LLMs), vector search, and modern AI technologies.

Responsibilities

  • Build AI-powered data acquisition, enrichment, and automation systems.
  • Design and develop scalable web crawlers, APIs, and data pipelines.
  • Build intelligent knowledge systems using LLMs, embeddings, vector databases, and AI agents.
  • Develop semantic search, recommendation, and discovery capabilities.
  • Own the end-to-end lifecycle of data collection, processing, validation, and indexing.
  • Build scalable crawlers and automated pipelines to collect data from websites, APIs, documentation, GitHub, research portals, directories, blogs, and other public sources.
  • Discover and maintain structured information on AI tools, models, datasets, research papers, benchmarks, APIs, MCP servers, companies, and emerging technologies.
  • Use LLMs to extract, summarize, classify, tag, and enrich structured and unstructured data.
  • Build systems for data validation, deduplication, entity resolution, and metadata enrichment.
  • Design semantic search, recommendation engines, vector search, and knowledge graph pipelines.
  • Develop backend services and APIs for data ingestion, processing, indexing, and retrieval.
  • Monitor crawler health, automate recurring data updates, and ensure data freshness and quality.
  • Integrate third-party APIs and external data sources while optimizing performance, scalability, and reliability.
  • Research and evaluate new AI technologies, models, and data sources to continuously improve platform intelligence.
  • Collaborate with Product and Engineering teams to build reliable, scalable, and AI-native data infrastructure.

Requirements

  • Strong proficiency in Python and backend development.
  • Experience with Playwright, Scrapy, Selenium, BeautifulSoup, or similar web crawling frameworks.
  • Experience building scalable ETL/data ingestion pipelines and REST APIs.
  • Strong understanding of HTML, JSON, APIs, web technologies, and data extraction techniques.
  • Experience with PostgreSQL, MongoDB, Redis, Elasticsearch/OpenSearch, or similar databases.
  • Familiarity with LLMs, AI APIs, embeddings, vector databases, Retrieval-Augmented Generation (RAG), and AI agents.
  • Experience with Git, Docker, and modern cloud development practices.
  • Strong analytical, debugging, and problem-solving skills.
  • Self-driven with the ability to work in a fast-paced startup environment and take ownership of end-to-end engineering problems.
  • Bonus: Experience with Knowledge Graphs, Hugging Face, GitHub APIs, LangChain, LlamaIndex, MCP ecosystem, Airflow, Temporal, or distributed systems.

More like this

Related opportunities