Tape transfer errors, ggus 1002113,

Tape Transfer Issues Report

Two separate issues contributed to the tape transfer failures observed over the past few weeks.

First, the space reservation for the tape buffer was being used in the SRR space usage report instead of the corresponding pool group. As a result, the space occupied by recalled files was not being accounted for correctly. This led Rucio to continue requesting additional files for transfer to tape under the false assumption that sufficient free space was still available. This issue has now been corrected.

Second, following the tape backend firmware upgrade, new problems appeared and some recalls became very slow. This created complications for the new bringonline+transfers model that NET2 started using at the request of DDM Ops. In cases where a recall completed only after the associated FTS transfer had already timed out or been cancelled, the cancellation was not propagated back to dCache. The recalled file therefore remained pinned for several days as an orphaned file. Since no transfer was available to move those files out of the tape buffer, the buffer filled rapidly. This issue is still under investigation, but we are now monitoring these events closely.

Unscheduled cluster downtime because the host certificate expired, and then because the certificate harvester uses to communicate with the cluster had expired.  This was followed by a brief blacklisting owing to a black hole node, which was removed from the cluster (it had developed a network issue during the downtime).

Short scheduled cluster downtime for OS upgrade on the main router.