IT-protoDUNE coordination (Single Phase and Double Phase)
Coordination Wiki: https://twiki.cern.ch/twiki/bin/viewauth/ProDUNEIT/WebHome
● Coordination update
Next meeting: 12th of July (Notice the change as during the meeting 11th was mentioned)
Attendants (identifiers to be used in minutes):
- Room
- Xavier Espinal (XE)
- Geoff Savage (GS)
- Andrea Dell'Acqua (AD)
- Ignacio Coterillo (IC)
- Stefan Lueders (SL)
- Cristian Contescu (CC)
- Remote:
- Maxim Potekhin (MP)
- Elisabetta Pennacchio (EP)
- Nektarios Benekos (NB)
- Stuart Fuess (SF)
Actions:
To provide a list of stakeholders which need to be contacted in case of interventions or outages on IT services of interest.
Note:
- Ignacio Coterillo (email: ignacio.coterillo.coz@cern.ch) will act as Secretary from now on. To be added to the watch list of related SNOW tickets in addition to Xavier Espinal
● shared/shifter accounts
- Question(GS): Is it possible to have a group account of an account that can be shared among shifters (similar to what is done in Fermilab)?
A (SL): CERN Service accounts would fit the description. Sharing passwords for service accounts is considered acceptable with the right means of sharing in place (e.g. using a encrypted file/management system)
- Q(GS): Is it possible to log into a service account using personal credentials, or is it required
A (SL): Technical aspects of implementations will vary per service use case. I suggest to put you in contact of some of the more technically/applied members of the Security Team.
- Q(GS): Is there any rule for Screensavers?
A (SL): They are not enforced, but its use is encouraged. For a 24/7 Control Room(CR) it should be OK to disable them assuming that other physical access and authentication controls are in place.
-- SL commented on a tentative implementation by CMS for separating RO/RW accesss to Control Room systems with the use of badges as an implementation to look up to.
- Q (AD): Is it possible to have stricter access control to the CR/Barracks?
A (SL): That access is in principle e-group controlled and at your discretion. Additional physical access control elements should be possible to install but the limitations there should be of time and money.
-- Remark by AD about how access control needs to be taking into account when planning training exercises
-- There was comment about whether to strongly restrict access to the CR/Barracks once the construction works have finished, but SL remarked access should be restricted since the moment when any changes that could be applied to controls accessible from the CR/Barracks may end up in critical damages
- Q (SL): Do you have a Safety responsible?
A (AD): Olga and Marcin? from Atlas
-- Remark (SL): regarding CCTV Cameras in CR needing to adhere to the right privacy regulations (e.g. Not continuous recording, online status visible, etc.)
● Round-table: status reports
- AD mentioned a request for additional EOS space (~50TBi)
- The space will be allocated under the Neutrino platform space but at a different level than the DUNE namespace. This is a different experiment (K2K related)
- DDQM Activities (MP)
- Successfully running a limited number of jobs. (No issues after switching from EOS Fuse to xrdcp. A realistic estimation of desired load should be the capacity to run ~300 concurrent jobs, but the infrastructure should be find with double that amount. Work is carried on rewriting some of the scripts and progress should be repoted within a couple of days
- Additional work for changing serveral of the Web pages
Q (AD): Would it be possible to have an Event Display in the CR at this time?
A (Maxim): Yes it should be possible but for now it would be displaying MC based data and as it's still under development subject to instabilities.
-- Maxim will provide related links
- Networking (Stuart Fuess-SF)
- Iptables issues due to a bad rule
- Running tests with EOS Public
Q (SF): Regarding the use of perSONAR for testing and measurement. Is there a perSONAR endpoint in the EOS realm?
A (XE): It is used regularly by WLCG and boxes are installed at the CC. FNAL as well should have a Perfsonar running, need to ensure dCache nodes are part of the OPN so perfSonar results make sense.
Remark (GS) still needs to receive all testing requirements and integrate them, as there seems to be several proposed methodologies which need to be put together for replicability. The general pipeline to be tested is:
Network -> Disk -> EOS
Remark (XE): The 10% guaranteed reservation on the 100GB link bandwidth will trigger only in case of saturation, doesn't imply a limit in case the link is not being fully used. The QoS settings have not been fully tested yet. For more information on this Eduardo Martelli from CS should be available for discussion.
- Elissabeta
- New servers have been installed
- Running a new data challenge to optimize the load of the local EOS instance
- Next step is to test FTS transfer, which should start next week
- EOS (CC): There will be a transparent intervention (failover) tomorrow in EOSPublic