Data Lakehouse Architecture (Banking, 1-year renewable contract)
evolution recruitment solutions pte. ltd.- Posted a day ago
- Be among the first 10 applicants
Job Description
Dear Applicant,
If you or someone you know is interested, please send the CV directly to [Confidential Information] (most preferred, as I may overlook some CVs due to the high volume).
Please note that visa sponsorship is not available at this time.
Key Responsibilities
- Own the end-to-end architecture and technical vision of an enterprise Lakehouse platform.
- Design and implement scalable data products, data marketplace, knowledge layers, and platforms supporting agentic workloads.
- Define target architectures for applications and platforms, with emphasis on reusability, scalability, performance, security, and operational efficiency.
- Develop and maintain technical roadmaps and architecture strategies for the Lakehouse platform.
- Establish technical frameworks and reusable patterns to accelerate the operationalisation of:
- Unstructured and multimodal content extraction.
- Lambda architecture and deployment patterns.
- Retrieval-Augmented Generation (RAG) and retrieval-augmented data patterns.
- Vector and graph-based data capabilities.
- Agentic workloads and AI-driven data solutions.
- Design and implement large-scale distributed and MPP compute workloads across on-premise, hybrid, and cloud environments.
- Architect and optimise Lakehouse platforms using open table formats, object storage, data federation, and multimodal query engines.
- Design hybrid and cloud architectures using private connectivity, workload placement strategies, Infrastructure-as-Code, and cloud cost optimisation.
- Design data contracts, SLAs, data quality rules, and governance standards for foundation and business data products in partnership with business stakeholders.
- Enable data products to be consumed by downstream applications through APIs, publish-subscribe mechanisms, generative BI, real-time dashboards, and data marketplaces.
- Support the architecture and implementation of RAG, embedding strategies, vector databases, graph databases, prompt engineering, and context management for agentic workloads.
- Provide technical quality assurance and ensure delivery conforms to defined software development methodologies, engineering standards, and technology practices.
- Review design specifications and technical deliverables produced by development teams.
- Create and maintain functional and non-functional specifications, architecture/design documents, deployment guides, and training materials.
- Independently install, customise, configure, and integrate software packages and technology solutions.
- Participate in RFPs, proof-of-concepts (POCs), and technology/product selection activities.
- Drive performance engineering, capacity planning, tuning, and optimisation of data platforms and workloads.
- Partner with business stakeholders, technology teams, vendors, and other technology functions to design and deliver enterprise solutions.
- Support continuous service improvement, process improvement, and operational excellence initiatives.
- Ensure effective integration with DevOps, CI/CD, monitoring, testing, and engineering toolchains.
- Provide technical guidance and mentorship while maintaining a high standard of quality across architecture and engineering deliveries.
Key Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
- 10-15 years of experience in Data Architecture, Big Data, Data Engineering, Data Lake, or Lakehouse implementations.
- Strong experience designing and implementing enterprise-scale Data Lakehouse platforms, preferably within the financial services industry.
- Hands-on experience with one or more major data/cloud platforms, such as Databricks, Snowflake, Cloudera, AWS, Azure, GCP, Huawei Cloud, or Alibaba Cloud.
- Strong experience with large-scale Lakehouse architecture, implementation, performance optimisation, and distributed computing.
- Deep knowledge of open table formats such as Apache Iceberg, Apache Hudi, and Delta Lake.
- Strong experience with object storage architecture, including hot, warm, and cold data tiering strategies.
- Experience with data federation technologies such as Trino, Denodo, and Dremio.
- Experience with multimodal/distributed query engines such as Hive, Impala, Apache Kudu, or equivalent technologies.
- Proven experience designing MPP and distributed compute workloads across on-premise, hybrid, and cloud environments.
- Strong understanding of hybrid and cloud architecture, including private connectivity technologies such as Direct Connect and ExpressRoute.
- Experience with workload placement, cloud architecture, egress cost optimisation, and Infrastructure-as-Code.
- Strong experience building and serving foundation and business data products through APIs, publish-subscribe/event-driven architectures, real-time dashboards, BI platforms, and data marketplaces.
- Experience supporting AI and agentic workloads, including:
- Retrieval-Augmented Generation (RAG).
- Embedding strategies.
- Vector databases.
- Graph databases.
- Prompt engineering.
- Context management.
- Agentic orchestration and workflows.
- Strong knowledge of modern vector database technologies, such as Databricks Vector Search, Azure AI Search, Pinecone, ChromaDB, Weaviate, or Snowflake Cortex.
- Knowledge of graph databases such as Neo4j, JanusGraph, TigerGraph, Cosmos DB, Amazon Neptune, or equivalent.
More Info
Key Skills
Alibaba Cloud
Data Management Services
Distributed Platforms
Open Table


