SRE/Infrastructure Engineer, ML Platform, Tokyo
- Tokyo
- Partial Remote
- Full-time
- August 10, 2026
Only applicants residing in Japan are eligible to apply.
Overview
Guided by Money Forward AI Vision 2026, Money Forward is driving company-wide AX (AI Transformation) to deliver "digital workers" — AI agents that carry out business operations autonomously. The CDAO Office leads the AI and data strategy that makes this possible across the entire group.
This is a dedicated SRE/infrastructure position within our MLOps Platform, focused specifically on the domain serving our Digital Bank and group-company fintech products — the shared credit evaluation model initiative known internally as "Financier".
Our MLOps Platform is more than a machine learning platform. It also encompasses the serving interfaces, authentication mechanisms, and security layers used by connected products, making it a group-wide product platform. The team currently operates as a scrum unit of a Product Manager, Engineering Manager, Data Scientists, and ML Engineers, running prototype validation. With the Digital Bank launch scheduled for FY2027, standing up infrastructure that can withstand high traffic and high availability requirements — along with production operations and monitoring — has become our most urgent priority.
A separate dedicated SRE owns the standard infrastructure for the MLOps Platform as a whole. In this role, you will work alongside that shared foundation while focusing your efforts on infrastructure built specifically for Digital Bank and fintech products. The scope may look broad at first glance, but it is deliberately bounded: by making full use of AI development tools such as Claude Code and by collaborating closely with our application and ML engineers, this is designed as a role that one dedicated engineer can own end to end.
Success in this role is clearly defined: ensuring the FY2027 Digital Bank launch succeeds from an infrastructure and reliability standpoint. Everything — platform build-out, security architecture, and observability — works backward from that milestone.
Mission Priorities
- Deliver the infrastructure and monitoring that keep serving interfaces and production model endpoints running reliably
- Design and continuously uphold financial-grade governance and security (mTLS, private network connectivity, and related controls)
- Once the above are in place, build out an MLOps environment that lets ML and application engineers develop, analyze, and deploy models efficiently
Responsibilities and Duties
As the primary SRE/infrastructure owner for the Digital Bank and fintech domain, you will:
- Build infrastructure and codify it with IaC (Terraform)
- Design, build, and continuously improve the AWS infrastructure behind our serving interfaces, data pipelines, and ML development platform, managed as code with Terraform.
- Design and monitor financial-grade network security
- Architect the infrastructure for mTLS (mutual TLS authentication), build and continuously improve private network connectivity for Digital Bank integration (AWS PrivateLink, VPC Peering, and similar), and expand security monitoring coverage.
- Improve observability and stand up operations and on-call
- Design and introduce monitoring for logs, metrics, and traces using CloudWatch and related tooling. Define SLOs and SLAs, establish an on-call rotation, and standardize incident response processes.
- Design and support the data platform / MLOps interface
- Design and build the interface between the Databricks platform operated by our DRE organization and the ML Ops Platform (data processing and pipeline integration), and provide infrastructure support on the ML Ops side.
- Drive infrastructure collaboration across the group
- Partner with the SRE who owns platform-wide ML Ops standards, as well as data platform and product platform teams in other organizations and group companies, to expand Digital Bank and fintech infrastructure while keeping it consistent with our cross-group foundation.
Required Skills and Experience
- Approximately 5+ years of professional experience as an SRE or infrastructure engineer
- Hands-on experience designing, building, and operating infrastructure on AWS using services such as ECS, API Gateway, VPC, and IAM
- Proven track record operating infrastructure using Terraform for infrastructure as code, module design, and CI/CD automation
- Solid grounding in advanced networking and security with experience designing architectures using VPC Peering, AWS PrivateLink, and TLS or mTLS authentication
- Experience building and operating observability systems including monitoring, alert design, and performance log analysis in cloud environments
- Working knowledge of CI/CD and container technologies with experience optimizing deployment pipelines using GitHub Actions and Docker
Preferred Skills and Experience
- Infrastructure support experience in MLOps or data platform domains using tools like Amazon SageMaker, Databricks, or Airflow
- Track record of standing up operations, defining SLOs and SLAs, and standardizing incident response using tools like Datadog and PagerDuty
- Security and audit experience in financial services or payments aligned with standards such as the FISC Security Guidelines
- Experience designing, building, and operating shared platforms in collaboration with data and product teams across multiple organizations
Language Requirements
- Business level Japanese (equivalent to JLPT N2 or above)
- Basic business level English (equivalent to TOEIC 700 or above)
- If you do not have a qualification equivalent to TOEIC 700 or above, you will be required to take a company-designated test during the selection process.
※Please note that the interviews in the selection process will be conducted in Japanese.
Who We’re Looking For
- Strong sense of ownership for production reliability and security to keep serving interfaces and private networking stable under high traffic capacity
- Drive to shape architecture and operational tooling during phases of genuine uncertainty by evaluating options and making key decisions
- Motivation to balance financial-grade security and governance with high developer productivity for application and machine learning teams
- Desire to support AI-era product platforms and embrace AI development tools to achieve outsized team impact
- Ability to solve problems autonomously and drive initiatives forward while collaborating closely with engineering teams and cross-company stakeholders
Technology Stack
- Infrastructure/Cloud: AWS (ECS Fargate, API Gateway, VPC, PrivateLink, SageMaker, S3, Glue, CloudWatch), Databricks
- DevOps/IaC/MLOps: Terraform, GitHub Actions, Docker
- Observability/Incident: CloudWatch
- AI Development Tools: Claude Code and similar
- Languages/Scripting: Python, Shell, SQL
- Communication/Project: Slack, Notion
※This list reflects what we have adopted to date. Observability SaaS and incident management tooling (Datadog, PagerDuty, and the like) are not yet in place — meaning you will have the opportunity to shape those technology decisions yourself.
Work Environment
At Money Forward, we provide an environment where we can create world-class services together, and we are looking forward to welcoming you.
- Provided PC Specs: We provide PCs equipped with the latest CPUs (MacOS or Windows). Custom-made PCs tailored to business requirements and replacements with the latest OS are also possible.
- Systems to Enhance the Development Environment: Peripheral devices necessary for work (such as displays, mice, keyboards) can be purchased as office supplies. Generally, you can choose from standard products (catalog), and if conditions are met, you can apply for non-standard products as well.
- Money Forward Library: We have a library system where you can freely borrow books, ranging from technical books to management books. Desired books can be purchased at the company's expense.
- Referral Driven: We cover the cost of recruitment meals. There is a referral reward system.
- Conference Participation Support: The company partially covers participation in domestic and international conferences, such as RubyKaigi and Google I/O.
About Money Forward
Money Forward, founded in 2012, strives to deliver exceptional value to users in various business domains. As a leading FinTech company, we offer over 40 services, ranging from personal finance management to B2B SaaS products.
We have been growing rapidly, and we are expanding our global hiring to help further expand the company. That means that we are open to hiring those with limited or no Japanese language proficiency.
Money Forward is one of Japan's hottest FinTech companies and it is now a great opportunity to be a part of one of our continued growths!
Get Job Alerts
Sign up for our newsletter to get hand-picked tech jobs in Japan – straight to your inbox.




