- Work model
- Remote
- Experience
- 7+ years
- Employment
- Not specified
- Compensation
- Not disclosed
- Technology signal
- 18 tags
Technology context
18Parsed from the vacancy text; ordered by relevance to this role.
AIAWSGenAIDatabasesMachine LearningPythonNLPReliability EngineeringScalabilityCI/CDAmazon OpenSearchAmazon S3AWS GlueAWS LambdaData Analytics EngineeringData Solution ArchitecturePySparkVector Databases
Full listing
Role description
We are seeking a Data Solution Architect, AWS, to lead data pipeline design and implementation for a cloud-native analytics and GenAI-enablement initiative, owning architecture decisions across data ingestion, transformation, storage, and search/retrieval infrastructure using AWS services.
Responsibilities
- Design and architect scalable data pipelines using AWS Glue for ETL orchestration, Lambda for event-driven processing, and S3 for data lake storage
- Define data architecture patterns aligned with AWS Well-Architected Framework principles for reliability, performance, and cost optimization
- Create technical specifications and architecture diagrams for data platform components
- Lead development of Python and PySpark solutions for large-scale data processing and transformation
- Architect OpenSearch and vector database solutions to support semantic search, retrieval-augmented generation, and AI/ML workloads
- Design data models and pipelines that enable advanced analytics including regression analysis and NLP applications
- Provide architectural guidance on GenAI strategy, advisory, and operational integration with existing data infrastructure
- Design infrastructure patterns that support machine learning model training, inference, and deployment workflows
- Advise on data preparation and feature engineering approaches for ML/AI use cases
- Establish CI/CD patterns for data pipeline deployment and infrastructure-as-code practices
- Collaborate with engineering teams to ensure data platform components meet operational excellence standards
- Support knowledge transfer and technical documentation for sustained delivery
Requirements
- 7+ years of experience in data analytics engineering
- High proficiency in PySpark
- Advanced expertise in AWS Glue, AWS Lambda, and Amazon S3
- Proven skills in Amazon OpenSearch
- English proficiency at B2 level or higher
Nice to have
- Knowledge of machine learning
- Familiarity with CI/CD
- Understanding of vector databases