DevOps Board
Target release | Ongoing |
|---|---|
Epic | DevOps Milestones |
Document status | ONGOING |
Document owner | @Gajendran C (Unlicensed) |
Objective
Ongoing DevOps epics and stories across various areas and enhancements around tools, infra and process.
Requirements
# | Requirement | User Story | Importance | Notes |
|
|---|---|---|---|---|---|
1 | Azure-as-an-additional | Azure playground setup with all the capabilities for a seamless option to choose b/w AWS or Azurehttps://digit-discuss.atlassian.net/browse/OPS-1 | SEVERE | - Deployment Manifest changes for Resources (S3, EBS, etc.) |
|
2 | GIT | Git Branching strategyhttps://digit-discuss.atlassian.net/browse/OPS-30 | SEVERE |
|
|
3 | Spinnaker | POC: Multicloud orchestration and deployment pipeline toolhttps://digit-discuss.atlassian.net/browse/OPS-31 | medium |
|
|
4 | Backup | Encrypted logs in S3https://digit-discuss.atlassian.net/browse/OPS-2 | HIGH |
|
|
5 | Backup | Logs backup for at least 6 months periodhttps://digit-discuss.atlassian.net/browse/OPS-3 | HIGH | - Need to do POC with CloudFront and Log Analytics before finding out our new Log life cycle solution |
|
6 | Capacity Planning | Cluster and app Sizing determinationhttps://digit-discuss.atlassian.net/browse/OPS-4 | medium | - Need to have sizing templates with BaseMin, GoodToHave & OptToHave |
|
7 | Infra | Node resizing/restructuring across all the env, including Punjab prod (Upon customer approval)https://digit-discuss.atlassian.net/browse/OPS-5 | backlog | - Need to change the instance type M4.large to M5.xLarge |
|
8 | Infra | Need to have dashboard, monitoring & Deployments for all the Envshttps://digit-discuss.atlassian.net/browse/OPS-6 | medium | - Need to evaluate a tool, which is cloud agnostic and all-in-one |
|
9 | Infra | backlog |
|
| |
10 | Kafka Improvement | Deploy HA Kafka and Zookeeper clusterhttps://digit-discuss.atlassian.net/browse/OPS-8 | sever | - Need to make Headless service configuration (Kafka connect) | - Requests go to hadrcoded individual nodes like Kafka0,1,2) |
11 | Kafka Improvement | Use Kakfa Connect to index instead of indexerhttps://digit-discuss.atlassian.net/browse/OPS-9 | sever |
|
|
12 | Kafka Improvment | Kakfa paritioning and multi consumer implementationhttps://digit-discuss.atlassian.net/browse/OPS-10 | severe |
| - 1-to-1 to 1-to-many |
13 | Kube Upgrade | Upgrade Kubernetes to 1.11.6 for all environments (Dev, QA, PUAT, PPROD)https://digit-discuss.atlassian.net/browse/OPS-11 | HIGH | - Kops upgrade | POC Done |
14 | Kube Upgrade | Pod auto-scaling strategyhttps://digit-discuss.atlassian.net/browse/OPS-12 | backlog |
|
|
15 | Logging | Move logging from direct ELK to ELK via Kafkahttps://digit-discuss.atlassian.net/browse/OPS-13 | severe |
| - Nithin is working on |
16 | Logging | Request/Response event logging from Zuulhttps://digit-discuss.atlassian.net/browse/OPS-14 | severe |
|
|
17 | Logging | Log maskinghttps://digit-discuss.atlassian.net/browse/OPS-15 | severe | - Need to get the List of Fields from Dev, before working on POC |
|
18 | Monitoring | Kafka Monitoring & Alertinghttps://digit-discuss.atlassian.net/browse/OPS-16 | HIGH |
|
|
19 | Monitoring | Move telemetry to internal Kafka and ELKhttps://digit-discuss.atlassian.net/browse/OPS-17 | severe |
| - Nithin is working on |
20 | Monitoring | Prometheus or any better monitoring, which is proactively reporting issues ahead of timehttps://digit-discuss.atlassian.net/browse/OPS-18 | HIGH |
|
|
21 | Monitoring | Zuul/NGINX Status code monitoring - New dashboardhttps://digit-discuss.atlassian.net/browse/OPS-19 | HIGH |
|
|
22 | Monitoring | Error monitoring and configurationhttps://digit-discuss.atlassian.net/browse/OPS-20 | HIGH |
|
|
23 | Monitoring | Health and Readiness check on all serviceshttps://digit-discuss.atlassian.net/browse/OPS-21 | severe |
|
|
24 | Monitoring | Intra service traffic management gatewayhttps://digit-discuss.atlassian.net/browse/OPS-22 | HIGH |
|
|
25 | Monitoring | Monitoring dashboards at multiple levels - infra, IT & businesshttps://digit-discuss.atlassian.net/browse/OPS-23 | HIGH |
|
|
26 | Process | IAM user policy for whole infrahttps://digit-discuss.atlassian.net/browse/OPS-24 | HIGH |
| - Only admins, team IAM users have access only for their respective S3 Buckets |
27 | RBAC | ACL on Kubectl access (after 1.11.6 upgrade)https://digit-discuss.atlassian.net/browse/OPS-25 | severe |
|
|
28 | RBAC | ACL in Jenkinshttps://digit-discuss.atlassian.net/browse/OPS-26 | severe |
|
|
29 | Release Mgmt | Need Helm like deployment strategy to rollout and rollback releases with one chart or single confighttps://digit-discuss.atlassian.net/browse/OPS-27 | medium |
|
|
30 | TBD | DB Masking and PII removalhttps://digit-discuss.atlassian.net/browse/OPS-28 | medium |
|
|
31 | Kube Encrypt | Modify kubernetes deployment encryptionhttps://digit-discuss.atlassian.net/browse/OPS-39 |
|
|
|
User interaction and design
Open Questions
Question | Answer | Date Answered |
|---|---|---|
|