Changelog
This changelog lists all updates, improvements and new features our Engineering team develops for our Skyscrapers Reference Developer Platform. These are rolled out automatically to all DevOps-as-a-Service customers.
2019 Q3
- 2019-08-29
Maintenance
Prometheus-blackbox-exporter available as optional cluster addon
We’ve added the prometheus-blackbox-exporter as a K8s cluster addon which can be enabled upon request. The blackbox exporter can be used for probing HTTP(S), DNS, TCP and ICMP endpoints, for example to check whether an external resource is up/down. …
- 2019-08-29
Maintenance
Concourse task that checks the status of the service after the deployment
We extended the functionallity for the ECS deployments with concourse. After the service gets deployed Concourse would just exit because Terraform doesn’t take the deployment itself into account. This resulted in having false deploys sometimes …
- 2019-08-28
Maintenance
Redshift monitoring via Prometheus
We have updated our stacks to support Redshift monitoring via the Prometheus Operator running on our K8s clusters. If you have Redshift running, you will now be able to see alerts in Alertmanager and on slack when there is something wrong with the cluster. …
- 2019-08-26
Maintenance
Neo4j monitoring via Prometheus
We have updated our stacks to support Neo4j monitoring (Neo4j >= 3.4) via the Prometheus Operator running on our K8s clusters. If you have Neo4j running, you will see metrics appearing in the new Neo4j Grafana dashboard. Previously we still monitored …
- 2019-08-20
Maintenance
Fix for dashboards HTTP 500 error when refreshing token
Since our SSO overhaul you might’ve been noticing sudden HTTP 500 errors while using the Alertmanager, Kubernetes of Prometheus dashboards when your token’s TTL expires. Normally when your OIDC token expires, your session should automatically …
- 2019-08-19
Maintenance
A note on CVE-2019-11247
Two weeks ago a patch for Kubernetes vulnerability CVE-2019-11247 was released for K8s 1.13, 1.14 and 1.15. Unfortunately as of writing clusters using older K8s versions (like our kops-based 1.11 clusters) are still vulnerable. In short this vulnerability …
- 2019-08-12
Maintenance
Kubernetes dashboards ERR_TOO_MANY_REDIRECTS bug
During the past days you might’ve been getting ERR_TOO_MANY_REDIRECTS and or Bad Request - Login session expired errors. This bug was introduced during last week’s cluster add-ons upgrade. We have reverted the change that’s causing these …
- 2019-08-09
Maintenance
Add Bitbucket, GitLab and Google authentication to Concourse
By default we only allowed authenticating to Concourse through GitHub and local users. It’s now possible to plug into other systems like Bitbucket, GitLab or Google. Let us know if you’d like to change to any of these authentication systems.
- 2019-08-06
Maintenance
Kubernetes add-on upgrades
In the following days we’ll be rolling-out a bunch of upgrades to the deployed add-ons on your clusters. You don’t have to do anything to apply these upgrades, we’ll do that for you. And it won’t cause any downtime to the cluster or …
- 2019-08-01
Maintenance
We're moving to EKS
The past months we’ve beeen heavily re-evaluting and testing AWS EKS as base for our reference solution. Today we can consider our platform GA and moving forward all new clusters will be setup using EKS. Naturally we’ll keep on supporting and …
- 2019-07-16
Maintenance
Cluster and Persistent Volume backups with Velero 1.0
Staging Kubernetes clusters are now backed up through Heptio Velero. Production rollout is happening in the following days. As default schedule, backups are taken each night (0:00 UTC) and are retained for 10 days, however these are configurable. Backups …
- 2019-07-09
Maintenance
SSO / OAuth2 overhaul
We’ve completely updated our cluster’s Single-Sign-On setup, adding new features and fixing some long-standing bugs. What has changed: DEX, which we use as a single Identity Service for all authentication within the cluster, has been separated …
2019 Q2
- 2019-06-06
Maintenance
Support for Cognito in ElasticSearch
in v2.3.8 we added support for Cognito and its options to our terraform-awselasticsearch module.
- 2019-06-04
Maintenance
Adding Prometheus monitoring for Elasticsearch on ECS
Our ECS monitoring solution now supports monitoring Elasticsearch clusters using Elasticsearch Exporter, Prometheus and AlertManager, so we can get notified via slack (critical/warnings) and via OpsGenie (critical) for any issues with ES. This is similar …
- 2019-04-17
Maintenance
Move to the AWS provided Kibana
We’re in the process of removing our kibana deployment from all the Staging clusters and replacing it with the AWS provided kibana setup that comes with the AWS ElasticSearch service. Production clusters will follow. This change will free up some …
- 2019-04-16
Maintenance
Update kube2iam to 0.10.7
We’ve updated kube2iam to the latest version (0.10.7) on all clusters. For context, kube2iam is the component that provides IAM credentials to containers running in your Kubernetes clusters without the need to distribute secrets. This new version of …
- 2019-04-10
Maintenance
Upgrade Concourse to version 5
During the comming days, we’ll roll out Concourse version 5.0.1 to all our setups. This is a major version upgrade, comming from version 4.2.3, and it includes some important new features and fixes. The most relevant change for Concourse users is …
- 2019-04-08
Maintenance
Increased monitoring alerts visibility
During the following days we’re going to rollout some changes in how Kubernetes monitoring notifications are delivered. From now on, all notifications comming from the production k8s monitoring system will be shown in our shared slack channel, that …
2019 Q1
- 2019-03-29
Maintenance
Upgrade to Kubernetes 1.11.9 [CVE-2019-1002100, CVE-2019-9946, CVE-2019-3874, CVE-2019-1002101]
We are in the process of upgrading our managed Kubernetes clusters from v1.11.6 to v1.11.9. Next to some general bugfixes and improvements, which you can find full details in the Kubernetes changelog, this rollout comes with several high and medium …
- 2019-03-19
Maintenance
Create simple AWS resources from K8s via the AWS Service Operator
We’ve made the AWS Service Operator available for deployment on our managed Kubernetes clusters. This Operator allows you to manage some AWS resources, like ECR repositrories and S3 buckets, by using Kubernetes Custom Resource Definitions. For …
- 2019-03-18
Maintenance
Support for cronjob monitoring
Update (18-03-2019): We found out there were enough default alerts covering all cases of cronjob failures. The following alerts are covering different failure cases accordingly: KubeJobCompletion: Warnning alert after 1 hour if any Job doesn’t …
- 2019-03-06
Maintenance
Upgrade Kubernetes components
We are in the process of upgrading our staging Kubernetes clusters components to the latest stable releases. Production clusters will follow in 1 to 2 weeks (to be announced) after we have confirmed there are no issues with our customer’s workloads. …
- 2019-03-06
Maintenance
Improved monitoring alerts on Slack
We have updated the format of the monitoring Slack notifications. You might have already noticed that the monitoring messages in your Slack channels now contain more useful information and are more structured. We’ve already started rolling out the …
- 2019-02-21
Maintenance
Mongodb monitoring and dashboards
We have updated the clusters to have support for mongodb monitoring, alerts and dashboards. If you have a mongodb cluster you will see that there is now a mongodb dashboard in Grafana and that we added specific alert rules for mongodb in prometheus.
- 2019-02-19
Maintenance
Improved etcd backups
We’ve upgraded all the k8s cluster with a new etcd backup implementation. The old backup solution was relying on daily snapshots taken from a service running in the master nodes. We’ve decided to take a new approach by using AWS Data Lifecycle …