Skip to content

Production Engineer

raccoonai

Taipei OfficeHybrid

Applying for this one?

We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.

Get my CV for this job

$25, one-time. No subscription.

【關於這個角色】

Raccoon AI 是新興的生成式 AI 客服解決方案提供商,專注協助電子商務與網路平台企業,自動化處理客服訊息。

我們的產品運行在正式 SaaS production environment,服務企業客戶並承諾 SLA。隨著客戶與系統規模持續成長,我們希望建立更成熟的 Production Engineering 機制,讓 production issue 能被快速定位、處理與預防,同時降低產品開發團隊頻繁被 incident 打斷的情況。

我們正在尋找一位 Production Engineer,負責 production issue 的第一線工程判斷、debugging 與修復,並持續改善整體服務穩定性。

【工作內容】

- 負責 SaaS production issue 的 engineering triage、debugging 與問題定位

- 根據 SLA 與 severity 判斷事件優先級,協助快速恢復服務

- 查看 application log、database、API request、queue、monitoring 等資訊,找出 root cause

- Reproduce 客戶回報的問題,判斷是產品 bug、資料問題、第三方服務異常或 infrastructure issue

- 能直接處理的 application bug,完成 hotfix、測試與 release

- 當問題需要深入 domain knowledge 時,整理完整技術資訊後 escalation 給對應 Feature Engineer

- 協助處理 Rails application、Python AI backend、API integration 與第三方服務相關問題

- 建立與維護 incident runbook、debugging tools、monitoring 與 alerting

- 分析 recurring incidents,找出系統性問題並推動改善

- 與 CS、PM、RD 協作,建立清楚的 incident response 與 escalation process

- 參與 postmortem,降低相同 production issue 再次發生的機率

【你需要具備】

- 3 年以上 Backend / Full-stack / Production Engineering 相關經驗

- 熟悉至少一種 Backend 技術棧,例如 Ruby on Rails、Python、Node.js 等

- 具備良好的 debugging 能力,能快速閱讀陌生 codebase 並定位問題

- 熟悉 SQL,能獨立進行 production database issue investigation

- 熟悉 REST API、Webhook、第三方 API integration 等常見 SaaS 架構

- 熟悉 log、monitoring、error tracking 等 production debugging 工具

- 理解 HTTP、network、database、queue、cache 等基本系統概念

- 能在資訊不完整的情況下快速整理問題、建立 hypothesis 並逐步排除

- 對 production reliability、incident handling 與 root cause analysis 有高度興趣

- 能與不同角色合作,清楚說明問題範圍、影響程度與處理進度

【加分條件】

- Ruby on Rails 或 Python production experience

- AWS / GCP / Kubernetes / Docker 經驗

- 熟悉 PostgreSQL、Redis、message queue

- 使用過 Datadog、Grafana、Sentry、CloudWatch 或類似 observability 工具

- 有 SaaS、B2B enterprise service 或 SLA environment 經驗

- 有 on-call、incident response、postmortem 經驗

- 有 AI / LLM application production experience

- 曾處理高流量 API、distributed system 或 third-party integration issue

【這個角色和一般 RD 有什麼不同】

一般 Product Engineer 的主要任務是開發新功能與產品 roadmap。Production Engineer 的主要任務則是確保已經上線的服務穩定運作,包括:

- 快速處理 production incident

- 縮短 MTTR

- 降低 SLA violation

- 找出 recurring issue 並消除 root cause

- 減少 Feature Engineer 被臨時 production issue 打斷

這不是單純「修 Bug」的角色,而是對 production system 有高度 ownership 的工程職位。

【我們重視的特質】

我們希望你看到 production issue 時,第一個反應不是追究這是誰寫的,而是思考「問題在哪裡、影響多少人、怎麼最快恢復、以及怎麼讓它不要再發生。」 如果你喜歡 troubleshooting、追 log、查 API、分析 database、理解整個系統如何運作,甚至覺得找出 tricky production bug 很有成就感,這個角色會非常適合你。

Seen 5 hours ago.

Original posting on raccoonai's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job