Overview
We’re seeking an experienced AI Data Architect to design and scale data/AI platforms that enable robust machine learning workflows and production-grade AI services. You will own the architecture for data pipelines, model-ready datasets, access controls, and retrieval-based generation (RAG) systems, working closely with ML engineers, data engineers, and security stakeholders.
Responsibilities
- Design end-to-end data architecture for AI/ML use cases, including ingestion, transformation, feature/data preparation, and dataset versioning
- Build and optimize scalable ML data pipelines on AWS to support training, evaluation, and inference workloads
- Own RAG architecture: data indexing, chunking strategies, embedding management, retrieval, and prompt/response integration patterns
- Define IAM roles, permissions, and secure access patterns for data, model artifacts, and vector stores
- Collaborate with engineers to implement data observability (lineage, quality checks, monitoring, and alerting) for AI pipelines
- Ensure reliability and performance by tuning storage formats, query patterns, and pipeline throughput
- Create architecture documentation and standards for reproducibility, scalability, and cost control
- Partner with stakeholders to translate business needs into technical design and measurable success criteria
Requirements
- 5+ years of experience designing data and/or AI platforms in production environments
- Strong hands-on experience with AWS (e.g., S3, IAM, and ML/data services) and building scalable data pipelines
- Solid understanding of AI/ML concepts and practical experience integrating ML workflows with production data systems
- Demonstrated experience with RAG systems, including embeddings/vector search and retrieval pipeline design
- Deep knowledge of IAM concepts and secure engineering practices (least privilege, role-based access, and auditability)
- Proficiency with data modeling and working with large-scale datasets
- Experience with deploying or supporting AI/ML services and handling data lifecycle (retention, access, and governance)
Nice to have
- Experience with vector databases or managed vector search solutions in AWS ecosystems
- Knowledge of LLM orchestration frameworks and production RAG evaluation methods
- Familiarity with data governance, PII handling, and compliance-oriented design
- Experience with CI/CD and infrastructure-as-code for ML/data deployments