About Cotomo
Cotomo is a Japan-based AI Character and AI Novel platform with more than 2 million users.
We are building a new form of consumer AI centered around characters, conversations, stories, and long-term relationships between humans and AI.
Our ambition is not limited to Japan. We are building Cotomo to become a globally competitive consumer AI platform.
Cotomo is a venture-backed startup supported by some of Japan’s leading venture capital firms, and we are entering the next stage of growth as we expand both our product and organization globally.
About the Role
We are looking for an ML Tech Lead to own Cotomo’s ML stack end-to-end — from experimentation and post-training to evaluation, inference, deployment, and infrastructure.
Our product uses a broad range of models across LLMs, ASR, TTS, image and video generation, and other multimodal systems. You will work across these technologies as the product evolves, rather than being tied to a single model or modality.
This is primarily an individual contributor role, not a traditional management position. You will spend most of your time building, experimenting, debugging, and shipping yourself, while also serving as the technical lead for our small ML team.
This is also not a pure research role. Our goal is to build the best possible product, not to build models in-house for their own sake.
Depending on the problem, the best solution may be an internally trained model, an open-weight model, a frontier API, or a combination of them. We expect you to make those decisions based on actual product quality, latency, reliability, and cost.
How We Work
Cotomo is a small, ambitious team operating with a high level of intensity and ownership.
Responsibilities are intentionally broad, and difficult problems are rarely confined to a single layer of the stack. In the same week, you might improve a conversational model, investigate a speech pipeline, integrate a new generative model, optimize GPU utilization, compare an external API with an in-house model, and work with Product on a new user experience.
This can be a demanding environment. It is best suited to people who enjoy high-intensity startup environments, substantial autonomy, and being personally responsible for important outcomes.
What You’ll Own
- Own Cotomo’s ML stack end-to-end across language, speech, image, video, and multimodal models
- Improve conversational quality, personality, memory, personalization, latency, reliability, and cost
- Build and improve training, post-training, evaluation, and experimentation pipelines
- Evaluate and A/B test frontier APIs, open-weight models, and in-house models, making pragmatic decisions based on product quality and unit economics
- Own production ML systems, including model serving, GPU infrastructure, observability, and cost optimization
- Work closely with Product and Engineering to turn new model capabilities into measurable user impact
- Provide technical direction, reviews, and hands-on support for our small ML team
What We’re Looking For
- 4+ years of full-time professional experience in machine learning, AI, or a closely related engineering role
- You’ve built and operated production ML systems and are comfortable owning them beyond the experimentation stage
- You have a strong understanding of modern deep learning and LLM systems across training, fine-tuning, evaluation, and inference
- Strong software engineering skills come naturally to you, including working with GPU-based ML infrastructure
- Ambiguous technical problems don’t need to be fully scoped before you start — you can define the problem, make trade-offs, and drive it through to production
- You think in terms of product outcomes, not just model benchmarks
- You’re pragmatic about architecture and are equally comfortable choosing an external API, an open-weight model, or an in-house solution when it is the best fit
- You can set technical direction and raise the bar for a small team while remaining primarily a hands-on individual contributor
Nice to Have
We’re especially interested if you have worked on:
- Conversational AI, AI companions, character AI, voice AI, or other consumer AI products
- ASR, TTS, image generation, video generation, or other multimodal systems
- Post-training, preference optimization, reinforcement learning, synthetic data, or personalization
- Distributed training or large-scale inference systems
- Latency, reliability, and cost optimization for production ML
- Staff, Principal, Lead, or senior Research/ML Engineer-level technical ownership
You’ll Thrive Here If
You enjoy staying close to the work and moving across boundaries when the problem requires it.
You’re likely to thrive at Cotomo if you prefer a small, high-ownership team over a heavily structured organization; if you’re comfortable moving between models, infrastructure, and product; and if the idea of building a globally competitive AI product from Japan feels exciting to you.
Role
- Title: ML Tech Lead
- Type: Primarily Individual Contributor
- Hands-on vs. Leadership: ~70–80% hands-on / ~20–30% technical leadership
- Scope: Language, speech and multimodal models, evaluation, inference, ML infrastructure, external APIs, and cost optimization
Language Requirements
- Japanese: Business level or Conversational Japanese?
- English: Fluent
Seniority
- Senior level or above
About Starley
Starley designs new relationships between people and AI by developing products that fit into everyday life.
They are developing "Cotomo", an audio-based AI conversation app, which users can use to talk about everyday life or even more personal topics.
Starley aims to redefine how AI and people communicate and to provide the world with new experiences.
Get Job Alerts
Sign up for our newsletter to get hand-picked tech jobs in Japan – straight to your inbox.








