Metrics startup fails with v3.10 because of missing corresponding images at the docker.io registry

Question

Metrics startup fails with v3.10 because of missing corresponding images at the docker.io registry

muehle28 opened this issue 6 years ago · comments

The cluster gets deployed as single-master/multi-nodes with the latest ansible playbook deploy_cluster (https://github.com/openshift/openshift-ansible/tree/release-3.10).

inventory file:
...
openshift_release=v3.10
openshift_image_tag=v3.10.0
...

oc get pods:
NAME READY STATUS RESTARTS AGE
hawkular-cassandra-1-l22q6 0/1 ImagePullBackOff 0 35m
hawkular-metrics-lt769 0/1 ImagePullBackOff 0 35m
hawkular-metrics-schema-sh789 0/1 ImagePullBackOff 0 37m
heapster-fshx6 0/1 ImagePullBackOff 0 35m

The following messages are extracts from the event-log for the according pods:

pulling image "docker.io/openshift/origin-metrics-cassandra:v3.10.0"
Failed to pull image "docker.io/openshift/origin-metrics-cassandra:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-metrics-cassandra:v3.10.0 not found

pulling image "docker.io/openshift/origin-metrics-hawkular-metrics:v3.10.0"
Failed to pull image "docker.io/openshift/origin-metrics-hawkular-metrics:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-metrics-hawkular-metrics:v3.10.0 not found

pulling image "docker.io/openshift/origin-metrics-schema-installer:v3.10.0"
Failed to pull image "docker.io/openshift/origin-metrics-schema-installer:v3.10.0": rpc error: code = Unknown desc = repository docker.io/openshift/origin-metrics-schema-installer not found: does not exist or no pull access

pulling image "docker.io/openshift/origin-metrics-heapster:v3.10.0"
Failed to pull image "docker.io/openshift/origin-metrics-heapster:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-metrics-heapster:v3.10.0 not found

It is possible to work around that for most of the images by specifying the release candidate tag with the following variables:

openshift_metrics_cassandra_image=docker.io/openshift/origin-metrics-cassandra:v3.10.0-rc.0
openshift_metrics_hawkular_metrics_image=docker.io/openshift/origin-metrics-hawkular-metrics:v3.10.0-rc.0
openshift_metrics_heapster_image=docker.io/openshift/origin-metrics-heapster:v3.10.0-rc.0

However, the origin-metrics-schema-installer is not available at all at the docker.io repository.

Michael Johann · Answer 1 · Fri Aug 03 2018 21:03:07 GMT+0800 (China Standard Time)

This is still the case after the release of 3.10

Colby Johnston · Answer 2 · Fri Aug 17 2018 00:42:56 GMT+0800 (China Standard Time)

Any update on this? The origin-metrics-schema-installer image is still not available.

aram.eth · Answer 3 · Sat Aug 18 2018 19:12:39 GMT+0800 (China Standard Time)

\cc @jsanda @smarterclayton

John Sanda · Answer 4 · Sat Aug 18 2018 22:22:12 GMT+0800 (China Standard Time)

Sorry for not responding sooner. @rubenvp8510 and I have been investigating and trying to figure out what happened. There are v3.10.0-rc.0 tags on docker hub, but no 3.10.0. This is not limited to just the schema installer image. I do not see a 3.10.0 tag for any of the images that are built in this repo.

aram.eth · Answer 5 · Tue Aug 21 2018 00:52:02 GMT+0800 (China Standard Time)

Any updates on v3.10.0 image tags?

Ivan Saldarriaga · Answer 6 · Tue Aug 21 2018 10:48:19 GMT+0800 (China Standard Time)

@jsanda

same error reported by a number of pods in the openshift-logging and openshift-node-problem-detector namespaces:

"docker.io/openshift/origin-logging-fluentd:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-logging-fluentd:v3.10.0 not found

Failed to pull image "docker.io/openshift/origin-logging-auth-proxy:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-logging-auth-proxy:v3.10.0 not found

Failed to pull image "docker.io/openshift/origin-logging-kibana:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/origin-logging-kibana:v3.10.0 not found

Failed to pull image "docker.io/openshift/node-problem-detector:v3.10.0": rpc error: code = Unknown desc = manifest for docker.io/openshift/node-problem-detector:v3.10.0 not found

could you share the docker build output?

Julien Francoz · Answer 7 · Wed Aug 22 2018 16:06:21 GMT+0800 (China Standard Time)

It's maybe related to this pending jenkins job : https://ci.openshift.redhat.com/jenkins/job/push_origin_metrics_release_310/ ?

Colby Johnston · Answer 8 · Thu Aug 30 2018 06:16:31 GMT+0800 (China Standard Time)

I was able to find a workaround for this issue, by using the following image from dockerhub.

docker.io/alv91/origin-metrics-schema-installer:latest

I updated the hawkular-metrics-schema pod with the above image and the job completed and hawkular and heapster pods then started.

Przemysław Robak · Answer 9 · Thu Aug 30 2018 15:53:13 GMT+0800 (China Standard Time)

@cjohnston23 I builded origin-metrics with this repo, switching to HAWKULAR_METRICS_VERSION="0.30.5.Final".

There is an issue, metrics are visible only on POD view. If you try to get metrics on eg. deplyment tab, you will get

Failed to perform operation due to an error: OnError while emitting onNext value: org.hawkular.metrics.model.Metric.class

EDIT:
I can confirm, that with openshift/origin-metrics-hawkular-metrics:v3.11 and alv91/origin-metrics-schema-installer:3.10 problem from above does not exist.

Jan-Otto Kröpke · Answer 10 · Sat Sep 01 2018 02:17:41 GMT+0800 (China Standard Time)

@jsanda Can you re-check this?

#429 (comment)

Hiren Vadalia · Answer 11 · Tue Sep 04 2018 14:55:36 GMT+0800 (China Standard Time)

Any update on this? I am also facing this issue while deploying OKD 3.10 with metrics.

Jan-Otto Kröpke · Answer 12 · Tue Sep 04 2018 15:39:11 GMT+0800 (China Standard Time)

My working solution:

# https://github.com/openshift/origin-metrics/issues/429
openshift_metrics_cassandra_image: "docker.io/openshift/origin-metrics-cassandra:v3.11.0"
openshift_metrics_hawkular_metrics_image: "docker.io/openshift/origin-metrics-hawkular-metrics:v3.11.0"
openshift_metrics_heapster_image: "docker.io/openshift/origin-metrics-heapster:v3.11.0"
# https://github.com/openshift/origin-metrics/issues/429#issuecomment-417124646
openshift_metrics_schema_installer_image: "docker.io/alv91/origin-metrics-schema-installer:v3.10.0"

Hiren Vadalia · Answer 13 · Tue Sep 04 2018 15:50:44 GMT+0800 (China Standard Time)

@jkroepke Thank a lots for confirming, I just see this comments and I am trying deployment with this changes in inventory

Konrad Mosoń · Answer 14 · Fri Sep 21 2018 20:22:26 GMT+0800 (China Standard Time)

This problem still exists, and @jkroepke solution works (but obviously this is different version… not-stable yet probably since 3.11 is not released yet?)

ThoTischner · Answer 15 · Wed Mar 20 2019 00:36:27 GMT+0800 (China Standard Time)

My working solution for upgrade path from 3.9 to 3.11:

ansible-playbook /usr/local/src/openshift-ansible/playbooks/openshift-metrics/config.yml -e ansible_ssh_user=foo -e openshift_image_tag=v3.10.0-rc.0 -K -T 60

OpenShift Bot · Answer 16 · Fri Sep 04 2020 12:05:23 GMT+0800 (China Standard Time)

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

OpenShift Bot · Answer 17 · Sun Oct 04 2020 13:58:16 GMT+0800 (China Standard Time)

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

OpenShift Bot · Answer 18 · Tue Nov 03 2020 15:43:50 GMT+0800 (China Standard Time)

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

OpenShift CI Robot · Answer 19 · Tue Nov 03 2020 15:44:08 GMT+0800 (China Standard Time)

@openshift-bot: Closing this issue.

In response to this:

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.