← Back to your search

Trend Micro

Staff/Sr. ML Infrastructure / Platform Engineer

Data & ML

Employment

Full-time
From the employer

Full time

Level

Staff / Senior

Category

Data & ML
Not doable from Austria

Country assessment

Not doable from Austria because it is tied to an office and never says it can be done remotely.

Our assessment is guidance. Confirm arrangements with the employer.

Skills mentioned in this posting

JavaScriptKubernetesTerraformAWSGCPLLMsCI/CDObservability

Job description

Join Trend ‧ Join New Generation

趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣 
===============================================================

About the Role 

We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models — from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling. 

 

Required Qualifications 

Model Serving & Inference 

  • Operate multi-model LLM serving infrastructure 
  • Tune autoscaling policies to balance GPU cost and latency SLAs 

Kubernetes & GPU Infrastructure 

  • Operate production K8s clusters with NVIDIA GPU nodes 
  • Handle GPU node lifecycle: NVIDIA driver setup 

Infrastructure as Code 

  • Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud 
  • Package platform components and model deployments as Helm charts 
  • Manage multi-environment configurations 

Observability & Performance 

  • Maintain monitoring stack: Prometheus, Grafana, 
  • Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token 
  • Set up alerting for SLA violations and OOM events 

 

Bonus Skills 

These are not required, but candidates with these skills will stand out. 

  • LoRA / PEFT fine-tuning workflows 
  • MLflow for experiment tracking, model registry, and automated adapter deployment 
  • Experience building LoRA adapter CI/CD pipelines (training → registry → serving) 
  • Experience with alternative inference frameworks such as SGLang or NVIDIA NIM, including deep Parameter Tuning for Continuous Batching, KV Cache management, and Speculative Decoding. 

===============================================================
連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World