Share this opportunity
Related Job
- DevOps Engineer, VNGGamesthành phố hồ chí minh
- CTV Kinh Doanh (Soundbox), Zalopaythành phố hồ chí minh
- IT Risk & Compliance Specialist, VNGGamesthành phố hồ chí minh
Job Search
Senior Software Engineer (Agentic Platform), GreenNode
OfficialTechSoftware26-ENG-4028
thành phố hồ chí minh
View this job in
English
Job description
About the role & product
We are building an **Agentic Platform** — a platform that helps customers deploy and operate AI agents at production scale, quickly and safely. The platform provides ready-made services / components so agent builders can integrate quickly instead of building from scratch — significantly shortening the time from idea to a finished agent. The platform also follows open protocols to easily connect with tools, data sources and other agents.
We believe a solid agentic platform is built on **foundational software engineering (backend & distributed systems)**, with AI / agents as the capability layer on top. Therefore, this role is designed with a **70% traditional software engineering / platform engineering** and **30% AI, AI agent and related protocols** split.
You will work directly with the engineering team to take the services powering the platform from architecture, implementation, to operation — ensuring high availability, low latency, security and scalability.
Main responsibilities
1. Platform & Backend Engineering (~70%)
- Design, implement and maintain backend services / APIs that meet high standards of **availability, latency and security**.
- Build the platform's infrastructure components to support the agent lifecycle — from where agents run, communication between agent / tool / service, state management, to security and authorization.
- Build a **secure sandbox runtime** where AI agents can execute generated code and tools in isolation — using containerization / microVMs (Docker, gVisor, Firecracker), namespace & seccomp isolation, resource limits (CPU / memory / network), and strict egress controls so untrusted agent actions never leak into the host or other tenants.
- Design data models and data flows on SQL (MySQL), NoSQL (MongoDB) and vector databases; ensure consistency, throughput and recoverability.
- Integrate message brokers / event-driven backbones (Kafka, RabbitMQ, AWS SQS/SNS) for asynchronous communication between the platform's internal services.
- Identify and resolve system issues: performance bottlenecks, memory leaks, race conditions, security vulnerabilities, resource leaks.
- Design and deploy **observability** systems (metrics, logs, distributed tracing) to monitor and debug the platform in production.
- Research, experiment with and evaluate new technology solutions to solve the system's technical problems.
2. AI & Agentic Capability (~30%)
- Design foundational services for AI agents; proactively keep up with trending AI agent features in the market so customers always access the latest capabilities.
- Build and integrate **MCP servers** as well as adapters for the **A2A** protocol to connect agents with tools, data sources and other agents.
- Integrate LLMs / agentic frameworks (LangChain, LangGraph, CrewAI, Strands…) into the platform in a framework-agnostic way; design abstractions so the platform does not depend on any specific model or framework.
- Design and build the **agent sandbox capability** — a code / tool execution environment (code interpreter, REPL, tool-call runtime) that gives AI agents the ability to run code, evaluate outputs and iterate safely; support multi-language execution, persistable session state, file I/O, and deterministic reproducibility so agents can reliably "think → execute → observe → refine".
- Define and enforce **sandbox security boundaries** — permission scoping per agent / per task (read-only fs, network allow-lists, approval gates for sensitive actions), and deep integration with the platform's authorization layer so sandbox actions are auditable, traceable and revocable.
- Write clients (SDK, CLI, AI agent skills) to help customers integrate and use the platform.
- Apply AI tooling (coding assistant, AI code review, AI-assisted testing & debugging) into the daily development and operations workflow to boost team productivity.
We are building an **Agentic Platform** — a platform that helps customers deploy and operate AI agents at production scale, quickly and safely. The platform provides ready-made services / components so agent builders can integrate quickly instead of building from scratch — significantly shortening the time from idea to a finished agent. The platform also follows open protocols to easily connect with tools, data sources and other agents.
We believe a solid agentic platform is built on **foundational software engineering (backend & distributed systems)**, with AI / agents as the capability layer on top. Therefore, this role is designed with a **70% traditional software engineering / platform engineering** and **30% AI, AI agent and related protocols** split.
You will work directly with the engineering team to take the services powering the platform from architecture, implementation, to operation — ensuring high availability, low latency, security and scalability.
Main responsibilities
1. Platform & Backend Engineering (~70%)
- Design, implement and maintain backend services / APIs that meet high standards of **availability, latency and security**.
- Build the platform's infrastructure components to support the agent lifecycle — from where agents run, communication between agent / tool / service, state management, to security and authorization.
- Build a **secure sandbox runtime** where AI agents can execute generated code and tools in isolation — using containerization / microVMs (Docker, gVisor, Firecracker), namespace & seccomp isolation, resource limits (CPU / memory / network), and strict egress controls so untrusted agent actions never leak into the host or other tenants.
- Design data models and data flows on SQL (MySQL), NoSQL (MongoDB) and vector databases; ensure consistency, throughput and recoverability.
- Integrate message brokers / event-driven backbones (Kafka, RabbitMQ, AWS SQS/SNS) for asynchronous communication between the platform's internal services.
- Identify and resolve system issues: performance bottlenecks, memory leaks, race conditions, security vulnerabilities, resource leaks.
- Design and deploy **observability** systems (metrics, logs, distributed tracing) to monitor and debug the platform in production.
- Research, experiment with and evaluate new technology solutions to solve the system's technical problems.
2. AI & Agentic Capability (~30%)
- Design foundational services for AI agents; proactively keep up with trending AI agent features in the market so customers always access the latest capabilities.
- Build and integrate **MCP servers** as well as adapters for the **A2A** protocol to connect agents with tools, data sources and other agents.
- Integrate LLMs / agentic frameworks (LangChain, LangGraph, CrewAI, Strands…) into the platform in a framework-agnostic way; design abstractions so the platform does not depend on any specific model or framework.
- Design and build the **agent sandbox capability** — a code / tool execution environment (code interpreter, REPL, tool-call runtime) that gives AI agents the ability to run code, evaluate outputs and iterate safely; support multi-language execution, persistable session state, file I/O, and deterministic reproducibility so agents can reliably "think → execute → observe → refine".
- Define and enforce **sandbox security boundaries** — permission scoping per agent / per task (read-only fs, network allow-lists, approval gates for sensitive actions), and deep integration with the platform's authorization layer so sandbox actions are auditable, traceable and revocable.
- Write clients (SDK, CLI, AI agent skills) to help customers integrate and use the platform.
- Apply AI tooling (coding assistant, AI code review, AI-assisted testing & debugging) into the daily development and operations workflow to boost team productivity.
3.Engineering Practice & Ownership (applies to both parts)
- Write clean code with unit and integration tests; maintain quality through code review.
- Own features / services from design to production; proactively propose architecture improvements.
- Mentorship and review for junior / mid engineers; contribute to building engineering culture.
- Work closely with product, infra and stakeholders to turn requirements into clear technical solutions.
- Write clean code with unit and integration tests; maintain quality through code review.
- Own features / services from design to production; proactively propose architecture improvements.
- Mentorship and review for junior / mid engineers; contribute to building engineering culture.
- Work closely with product, infra and stakeholders to turn requirements into clear technical solutions.
Requirement
Job requirements
I/ Must-have
1. Background & experience
- Bachelor's degree or above in Computer Science, Engineering or a related field.
- Experience building backend services / distributed systems running in production; for us, years of experience is only a reference number — **actual capability** is the deciding factor.
- Solid foundation in **software architecture, design patterns, distributed systems** and best practices.
- Strong problem-solving, logical thinking and analytical skills; effective communication and collaboration.
2. Backend & Data
- Proficient in at least one of: **Java (Spring Boot), Go, Python**.
- Proficient in **SQL (MySQL)** and **NoSQL (MongoDB)**; understand indexing, query optimization, transactions and data modeling.
- Hands-on experience with message brokers: **Kafka, RabbitMQ, ActiveMQ** or **AWS SQS/SNS**.
- Good understanding of **REST API** design, authentication / authorization (OAuth, API key, service-to-service auth) and security fundamentals.
3.Infra & DevOps
- Familiar with **Git, Docker, CI/CD**; understand microservices deployment.
- Understand **observability** (metrics, logs, tracing) and have worked with related tools (Prometheus/Grafana, ELK/OpenSearch, OpenTelemetry, Jaeger…).
4. AI & Agentic
- Understand the core concepts of agent architecture: **tool use, memory/context, orchestration loop, guardrails, multi-agent coordination**.
- Understand and have worked with **MCP (Model Context Protocol)** — **knowing how to write an MCP server is a big plus**.
- Understand **A2A (Agent-to-Agent)** and other agent protocols (ADK, ReAct).
- Have used agentic frameworks (LangChain, LangGraph, CrewAI…) and/or RAG / vector databases.
- Familiar with using AI tooling in daily work.
II/ Nice-to-have
- Have built or contributed to a real **agentic platform / AI platform**.
- Have written an **MCP server** and integrated it into a production system.
- Have deployed an end-to-end **monitoring / observability** system.
- Experience with **Kubernetes** and operating container workloads at scale.
- Experience with **code interpreter sandbox**, identity / authorization for agents, or semantic caching.
- Have built a **secure sandbox / code execution runtime** for AI agents or multi-tenant workloads — hands-on with container isolation (Docker, gVisor, Firecracker / microVMs), namespace & seccomp, network policies, or serverless code execution platforms.
- **Good English listening and speaking skills** — a plus for working with international teams, partners and technical documentation.
I/ Must-have
1. Background & experience
- Bachelor's degree or above in Computer Science, Engineering or a related field.
- Experience building backend services / distributed systems running in production; for us, years of experience is only a reference number — **actual capability** is the deciding factor.
- Solid foundation in **software architecture, design patterns, distributed systems** and best practices.
- Strong problem-solving, logical thinking and analytical skills; effective communication and collaboration.
2. Backend & Data
- Proficient in at least one of: **Java (Spring Boot), Go, Python**.
- Proficient in **SQL (MySQL)** and **NoSQL (MongoDB)**; understand indexing, query optimization, transactions and data modeling.
- Hands-on experience with message brokers: **Kafka, RabbitMQ, ActiveMQ** or **AWS SQS/SNS**.
- Good understanding of **REST API** design, authentication / authorization (OAuth, API key, service-to-service auth) and security fundamentals.
3.Infra & DevOps
- Familiar with **Git, Docker, CI/CD**; understand microservices deployment.
- Understand **observability** (metrics, logs, tracing) and have worked with related tools (Prometheus/Grafana, ELK/OpenSearch, OpenTelemetry, Jaeger…).
4. AI & Agentic
- Understand the core concepts of agent architecture: **tool use, memory/context, orchestration loop, guardrails, multi-agent coordination**.
- Understand and have worked with **MCP (Model Context Protocol)** — **knowing how to write an MCP server is a big plus**.
- Understand **A2A (Agent-to-Agent)** and other agent protocols (ADK, ReAct).
- Have used agentic frameworks (LangChain, LangGraph, CrewAI…) and/or RAG / vector databases.
- Familiar with using AI tooling in daily work.
II/ Nice-to-have
- Have built or contributed to a real **agentic platform / AI platform**.
- Have written an **MCP server** and integrated it into a production system.
- Have deployed an end-to-end **monitoring / observability** system.
- Experience with **Kubernetes** and operating container workloads at scale.
- Experience with **code interpreter sandbox**, identity / authorization for agents, or semantic caching.
- Have built a **secure sandbox / code execution runtime** for AI agents or multi-tenant workloads — hands-on with container isolation (Docker, gVisor, Firecracker / microVMs), namespace & seccomp, network policies, or serverless code execution platforms.
- **Good English listening and speaking skills** — a plus for working with international teams, partners and technical documentation.
Registration was successful!
We've received your profile and we do appreciate your interest in our job opportunities. We will screen your application and contact you for further steps if you are short-listed. Otherwise, the application with no response received within 2 weeks is considered unsuitable application, and we will keep your resume in our database and may consider for appropriate future openings. Again, thank you for considering VNG as a potential employer.
