HOT JOBS

OPENING SEARCH

Your location::Home>HOT JOBS>USA>
发布日期:
2016-03-14 15:18:55
职位名称:
Cloud Operation Tools Architect
所属行业:
Telecom/Wireless
所在省份:
USA
薪资范围:
联系方式:
roger.zuo@dsi.com.cn

职位描述

A Global ICT vendor from China....

Position Responsibility:

This role is responsible to design a highly reliable and scalable alerting and incident resolution platform to effectively support The Client’s public cloud and its partners’ cloud operation. In particular, he/she should:
  • Lead/contribute to design a highly effective alerting platform smoothly integrated with The Client’s next generation monitoring platform. .
  • Architect the end to end incident management infrastructure for The Client Public Cloud, including , effective event aggregation and classification, alarm suppression, oncall scheduling, notification, and automated escalation mechanisms and so on.
  • Guide and participate external technology communication and cooperation, including enterprise customers, IT vendors, ISP/ICP, university, etc. Establish the channels with the best experts in IT industry. Identify and catch the opportunity on the technology purchasing, collaboration and innovation

2.2 Position requirements
Basic Requirements Education background   B.Sc./M.Sc.in computer science, engineering science, software engineering or similar relevant discipline
Work experience 5+ years working experience in building alarming system for large scale public cloud
Language (optional) English can be used as working language
Professional Knowledge Professional knowledge  
•Familiar with large-scale cloud alerting system development and operation
Key Skillset Professional skills
  • Deep understanding in the typical alarming system for operating an online service in AWS or Azure. Hands on experience in designing at least one customer facing online service.
  • Deep understanding in online service operation challenges: have opinions about what works and what doesn’t; understand the importance to distinguish alarms from noises; can accurately define the severity of incidents for a proactive issue resolution etc.
  •  Excellent knowledge of concepts like consistency, availability, real-time dispatching, and distributed queuing. Know how to make the appropriate tradeoffs when designing a system, and that the “right choice” will often differ based on requirements, not on dogma. .
  • Enjoy being challenged by problems of scale and complexity, with a strong desire to make services better for users.
  • Has broad knowledge on IT industry, especially on cloud computing, virtualization, storage, networking, middleware, database, web architecture etc.