Opportunity brief
Ascendion is hiring for the role of AI+Data Engineer!
Responsibilities of the candidate:
- Hands-on experience with Apache Spark or PySpark. - Experience with cloud platforms such as Azure, AWS, or GCP. - Knowledge of data storage technologies, including Data Lakes, Data Warehouses, and Delta Lake. - Familiarity with orchestration tools such as Apache Airflow or Azure Data Factory. - Understanding of machine learning workflows and MLOps concepts. - Experience with Git, CI/CD pipelines, and Agile methodologies. - Strong analytical and problem-solving skills. - Experience with Databricks is preferred. - Knowledge of Generative AI, LLMs, LangChain, Azure OpenAI, or similar AI frameworks is preferred. - Familiarity with vector databases such as Pinecone, FAISS, or ChromaDB. - Experience with Docker and Kubernetes. - Knowledge of streaming technologies such as Kafka or Azure Event Hubs. - Understanding of data governance and security best practices.
Requirements:
- Design, develop, and maintain scalable ETL/ELT pipelines for structured and unstructured data. - Build and optimize data pipelines for AI/ML and Generative AI applications. - Prepare, clean, transform, and validate datasets for machine learning model training and inference. - Integrate AI/ML models into production environments and applications. - Develop and maintain data lakes, data warehouses, and feature stores. - Work with cloud platforms such as Azure, AWS, or GCP to deploy and manage data solutions. - Implement data quality, governance, security, and monitoring practices. - Optimize SQL queries and big data processing jobs for performance, scalability, and efficiency. - Collaborate with Data Scientists, ML Engineers, and Software Engineers to deliver AI-powered products and solutions. - Create and maintain technical documentation for data and AI pipelines. - Support CI/CD processes for data engineering and AI/ML pipelines. - Strong proficiency in Python and SQL. - Experience with ETL/ELT tools and data pipeline development.