DevOps Board

DevOps Board

Target release

Ongoing

Epic

DevOps Milestones 

Document status

ONGOING

Document owner

@Gajendran C (Unlicensed)

Objective

Ongoing DevOps epics and stories across various areas and enhancements around tools, infra and process. 

Requirements

#

Requirement

User Story

Importance

Notes

 

1

Azure-as-an-additional

Azure playground setup with all the capabilities for a seamless option to choose b/w AWS or Azurehttps://digit-discuss.atlassian.net/browse/OPS-1

SEVERE

- Deployment Manifest changes for Resources (S3, EBS, etc.)
- Eng support: application level changes from S3 to Azure Blob for FileStore, Telemetry, Logos

  • Playground Ready (POC Env)

  • Data migration from AWS to Azure Tested

  • All services are intact with PBprod

2

GIT 

Git Branching strategyhttps://digit-discuss.atlassian.net/browse/OPS-30

SEVERE

  • Git Review process is implemented

  • Branching and multirepo strategy is in progress

 

3

Spinnaker

POC: Multicloud orchestration and deployment pipeline toolhttps://digit-discuss.atlassian.net/browse/OPS-31

medium

  • Being able to orchestrate and deploy across multicluster and multi cloud platform.

  • Being able to create RBAC and assign deployment pipelines. 

 

4

Backup

Encrypted logs in S3https://digit-discuss.atlassian.net/browse/OPS-2

HIGH

 

 

5

Backup

Logs backup for at least 6 months periodhttps://digit-discuss.atlassian.net/browse/OPS-3

HIGH

- Need to do POC with CloudFront and Log Analytics before finding out our new Log life cycle solution

 

6

Capacity Planning

Cluster and app Sizing determinationhttps://digit-discuss.atlassian.net/browse/OPS-4

medium

- Need to have sizing templates with BaseMin, GoodToHave & OptToHave

 

7

Infra

Node resizing/restructuring across all the env, including Punjab prod (Upon customer approval)https://digit-discuss.atlassian.net/browse/OPS-5

backlog

- Need to change the instance type M4.large to M5.xLarge

 

8

Infra

Need to have dashboard, monitoring & Deployments for all the Envshttps://digit-discuss.atlassian.net/browse/OPS-6

medium

- Need to evaluate a tool, which is cloud agnostic and all-in-one

 

9

Infra

MultiCloudhttps://digit-discuss.atlassian.net/browse/OPS-7

backlog

 

 

10

Kafka Improvement

Deploy HA Kafka and Zookeeper clusterhttps://digit-discuss.atlassian.net/browse/OPS-8

sever

- Need to make Headless service configuration (Kafka connect)

- Requests go to hadrcoded individual nodes like Kafka0,1,2)

11

Kafka Improvement

Use Kakfa Connect to index instead of indexerhttps://digit-discuss.atlassian.net/browse/OPS-9

sever

 

 

12

Kafka Improvment

Kakfa paritioning and multi consumer implementationhttps://digit-discuss.atlassian.net/browse/OPS-10

severe

 

- 1-to-1 to 1-to-many

13

Kube Upgrade

Upgrade Kubernetes to 1.11.6 for all environments (Dev, QA, PUAT, PPROD)https://digit-discuss.atlassian.net/browse/OPS-11

HIGH

- Kops upgrade
- Manifests changes

POC Done

14

Kube Upgrade

Pod auto-scaling strategyhttps://digit-discuss.atlassian.net/browse/OPS-12

backlog

 

 

15

Logging

Move logging from direct ELK to ELK via Kafkahttps://digit-discuss.atlassian.net/browse/OPS-13

severe

 

- Nithin is working on

16

Logging

Request/Response event logging from Zuulhttps://digit-discuss.atlassian.net/browse/OPS-14

severe

 

 

17

Logging

Log maskinghttps://digit-discuss.atlassian.net/browse/OPS-15

severe

- Need to get the List of Fields from Dev, before working on POC

 

18

Monitoring

Kafka Monitoring & Alertinghttps://digit-discuss.atlassian.net/browse/OPS-16

HIGH

 

 

19

Monitoring

Move telemetry to internal Kafka and ELKhttps://digit-discuss.atlassian.net/browse/OPS-17

severe

 

- Nithin is working on

20

Monitoring

Prometheus or any better monitoring, which is proactively reporting issues ahead of timehttps://digit-discuss.atlassian.net/browse/OPS-18

HIGH

 

 

21

Monitoring

Zuul/NGINX Status code monitoring - New dashboardhttps://digit-discuss.atlassian.net/browse/OPS-19

HIGH

 

 

22

Monitoring

Error monitoring and configurationhttps://digit-discuss.atlassian.net/browse/OPS-20

HIGH

 

 

23

Monitoring

Health and Readiness check on all serviceshttps://digit-discuss.atlassian.net/browse/OPS-21

severe

 

 

24

Monitoring

Intra service traffic management gatewayhttps://digit-discuss.atlassian.net/browse/OPS-22

HIGH

 

 

25

Monitoring

Monitoring dashboards at multiple levels - infra, IT & businesshttps://digit-discuss.atlassian.net/browse/OPS-23

HIGH

 

 

26

Process

IAM user policy for whole infrahttps://digit-discuss.atlassian.net/browse/OPS-24

HIGH

 

- Only admins, team IAM users have access only for their respective S3 Buckets

27

RBAC

ACL on Kubectl access (after 1.11.6 upgrade)https://digit-discuss.atlassian.net/browse/OPS-25

severe

 

 

28

RBAC

ACL in Jenkinshttps://digit-discuss.atlassian.net/browse/OPS-26

severe

 

 

29

Release Mgmt

Need Helm like deployment strategy to rollout and rollback releases with one chart or single confighttps://digit-discuss.atlassian.net/browse/OPS-27

medium

 

 

30

TBD

DB Masking and PII removalhttps://digit-discuss.atlassian.net/browse/OPS-28

medium

 

 

31

Kube Encrypt

Modify kubernetes deployment encryptionhttps://digit-discuss.atlassian.net/browse/OPS-39 

 

 

 

User interaction and design

Open Questions

Question

Answer

Date Answered

 

Out of Scope