Color code: (critical, news during the meeting: green, news from this week: blue, news from last week: purple, no news: black)
High priority framework topics:
- Problem at end of run with lots of error messages and breaks many calibration runs https://alice.its.cern.ch/jira/browse/O2-3315
- Unfortunately, online there are still errors after fixing those that are seen in the FST (e.g. partition 2e5XVwXYwvX).
- After a recent O2 update, these messages started to appear in the FST as well again. At least this case can be reproduced locally.
- Will continue debugging with Giulio next week
- Fix START-STOP-START for good
- Multi-threaded pipeline still not working in FST / sync processing, but only in standalone benchmark.
- Discussed with Giulio that it might not converge until Pb-Pb, so I'll try to implement a standalone multi-threading solution in DPL.
- Problem with QC topologies with expendable tasks- For items to do see: https://alice.its.cern.ch/jira/browse/QC-953 - Status?
- Problem in QC where we collect messages in memory while the run is stopped: https://alice.its.cern.ch/jira/browse/O2-3691
- Had to be reverted, since it broke QC. Proper fix needs at least a new FMQ version, (PR by Giulio is open), and perhaps new work. Will merge it step by step in alidist / O2, but probably not going to be deployed at P2 until after Pb-Pb, if QC is not strongly pushing for it.
Other framework tickets:
- TOF problem with receiving condition in tof-compressor: https://alice.its.cern.ch/jira/browse/O2-3681
- Grafana metrics: Might want to introduce additional rate metrics that subtract the header overhead to have the pure payload: low priority.
- Backpressure reporting when there is only 1 input channel: no progress.
- Stop entire workflow if one process segfaults / exits unexpectedly. Tested again in January, still not working despite some fixes. https://alice.its.cern.ch/jira/browse/O2-2710
- https://alice.its.cern.ch/jira/browse/O2-1900 : FIX in PR, but has side effects which must also be fixed.
- https://alice.its.cern.ch/jira/browse/O2-2213 : Cannot override debug severity for tpc-tracker
- https://alice.its.cern.ch/jira/browse/O2-2209 : Improve DebugGUI information
- https://alice.its.cern.ch/jira/browse/O2-2140 : Better error message (or a message at all) when input missing
- https://alice.its.cern.ch/jira/browse/O2-2361 : Problem with 2 devices of the same name
- https://alice.its.cern.ch/jira/browse/O2-2300 : Usage of valgrind in external terminal: The testcase is currently causing a segfault, which is an unrelated problem and must be fixed first. Reproduced and investigated by Giulio.
- Found a reproducible crash (while fixing the memory leak) in the TOF compressed-decoder at workflow termination, if the wrong topology is running. Not critical, since it is only at the termination, and the fix of the topology avoids it in any case. But we should still understand and fix the crash itself. A reproducer is available.
- Support in DPL GUI to send individual START and STOP commands.
- Problem I mentioned last time with non-critical QC tasks and DPL CCDB fetcher is real. Will need some extra work to solve it. Otherwise non-critical QC tasks will stall the DPL chain when they fail.
Global calibration topics:
- TPC IDC workflow problem - no progress.
- TPC has issues with SAC workflow. Need to understand if this is the known long-standing DPL issue with "Dropping lifetime::timeframe" or something else.
- Even with latest changes, difficult to ensure guaranteed calibration finalization at end of global run (as discussed with Ruben yesterday).
- After discussion with Peter, Giulio this morning: We should push for 2 additional states in the state machine at the end of run between RUNNING and READY:
- DRAIN: For all but O2, the transision RUNNING --> FINALIZE is identical to what we do in STOP: RUNNING --> READY at the moment. I.e. no more data will come in then.
- O2 could finalize the current TF processing with some time out, where it stops processing incoming data, and at EndOfStream trigger the calibration postprocessing.
- FINALIZE: No more data is guaranteed to come in, but the calibration could still be running. So we leave FMQ channels open, and have a time out to finalize the calibration. If input proxies have not yet received the EndOfStream, they will inject it to trigger the final calibration.
- This would require changes in O2, DD, ECS, FMQ, but all changes except for in O2 should be trivial, since all other components would not do anything in these states.
- Started to draft a document, but want to double-check it will work out this way before making it public.
- output-proxy cannot send the same data to 2 calib workflows.
- Fixed after debug session by Giulio and David.
- First version of the fix had a performance regression, breaking FLP processing, which was afterwards fixed by Giulio.
- EMCAL bad channel calib too slow for high rate pb-pb.
- EMCAL has now a multi-threaded version, and also the option to downscale, steered via extra env variables. Should be ok for Pb-Pb.
- Problem when calibration aggregators suddenly receive an endOfStream in the middle of the run and stop processing:
- Happens since ODC 0.78, which checks device states during running: If one device fails, it sends sigkill to all devices of the collection, and FairMQ takes the shortest way through the state machine to EXIT, which involves a STOP transition, which then sends the EoS. The DPL input proxy on the calib node should in principle check that it has received EoS from all nodes, but for some reason that is not working. Todo:
- Fix the input proxy to correctly count the number of EoS.
- Change the device behavior such that we do not send EoS in case of sigkill when running on FLP/EPN.
CCDB:
- Performance issues accessing https://alice-ccdb.cern.ch by Matteo and me.
- Understood and fixed: JAliEn-ROOT doesn't have a developer right now, so was looking myself.
- It uses libwebsockets in an incorrect way, driving the event loop from the main thread, without an event loop library like libuv, and without spawning an extra thread.
- When nothing to do, libwebsocket goes to sleep, until it is woken up by either a signal from the OS, or from activity on a socket.
- With never versions, libwebsocket optimized the overhead, so got woken up less often, thus making the delays longer for us.
- Essentially, CCDB access was slow, since the processes were just sleeping, until woken up by the OS randomly.
- Current fix in libwebsocket works around that in the current implementation, by cancelling the sleep after checking the connection state. In principle, we should update JAliEn-ROOT to a modern libwebsocket usage supporting an external event loop.
- Propose to merge this, and also bump libwebsockets in to a more recent version (current is from 2018 and unmaintained).
Sync reconstruction:
- Asked detectors to modernize their standalone workflows, and follow several design patterns of global workflow:
- Use helper to fetch and catch QC JSONs from consul.
- Check for SYNTHETIC / STAGING run, and do not upload CCDB objects to production CCDB in this case.
- Support $GLOBALDPLOPT, needed for MI100 workflows.
- Use the commom helpers we provide to assemble the merged workflow string.
- Most detectors have implemented the requested changes meanwhile.
Async reconstruction
- Remaining oscilation problem: GPUs get sometimes stalled for a long time up to 2 minutes. No progress
- Checking 2 things: does the situation get better without GPU monitoring? --> Inconclusive
- We can use increased GPU processes priority as a mitigation, but doesn't fully fix the issue.
- EPN vobox for MI100:
- voboxes recreated and moved to new IP range.
- Routing of second gateway fixed, vobox2 updated to use that as default route.
- Fixed vobox DNS entries, after EPN changed DNS entries for gateway nodes, which again caused issues for MonaLisa.
- Performance issue seen in async reco on MI100, need to investigate.
EPN major topics:
- Fast movement of nodes between async / online without EPN expert intervention.
- 2 goals I would like to set for the final solution:
- It should not be needed to stop the SLURM schedulers when moving nodes, there should be no limitation for ongoing runs at P2 and ongoing async jobs.
- We must not lose which nodes are marked as bad while moving.
- Interface to change SHM memory sizes when no run is ongoing. Otherwise we cannot tune the workflow for both Pb-Pb and pp: https://alice.its.cern.ch/jira/browse/EPN-250
- Lubos to provide interface to querry current EPN SHM settings - ETA July 2023, Status?
- Improve DataDistribution file replay performance, currently cannot do faster than 0.8 Hz, cannot test MI100 EPN in Pb-Pb at nominal rate, and cannot test pp workflow for 100 EPNs in FST since DD injects TFs too slowly. https://alice.its.cern.ch/jira/browse/EPN-244 NO ETA
- Interface to communicate list of active EPNs to epn2eos monitoring: https://alice.its.cern.ch/jira/browse/EPN-381
- Alice enabled the new mails yesterday. For now I subscribed to it, to see how it goes. If the mail-flood is avoided now, we can subscribe more people.
- DataDistribution distributes data round-robin in absense of backpressure, but it would be better to do it based on buffer utilization, and give more data to MI100 nodes. Now, we are driving the MI50 nodes at 100% capacity with backpressure, and then only backpressured TFs go on MI100 nodes. This increases the memory pressure on the MI50 nodes, which is anyway a critical point. https://alice.its.cern.ch/jira/browse/EPN-397
Other EPN topics:
EPN farm upgrade:
- Finally managed to obtain a working ROCm setup that runs on MI50 and MI100 (without workarounds in our code, and without performance impact):
- ROCm 5.5.3 basis.
- ROCm 5.6 kernel module (cannot use 5.6 due to new regression in compiler)
- Custom 5.5.1 compiler with a commit revert (revert is a workaround, AMD is still checking what exactly goes wrong).
- Repeated high rate tests with larger number of EPNs (~330 nodes used)
TPC Raw decoding checks:
- Add additional check on DPL level, to make sure firstOrbit received from all detectors is identical, when creating the TimeFrame first orbit.
Full system test issues:
Topology generation:
- Should test to deploy topology with DPL driver, to have the remote GUI available. Status?
Software deployment at P2.
- No software update this week yet, tentatively on Friday.
- Most likely last regular update, afterwards only cherry-picks of improvements / bug fixes requested by RC.
- We do only cherry-picks that apply cleanly, i.e. if a new feature cannot be cherry-picked, we would either bump to latest dev or not deploy it.
- New CTF coding was not properly validated yet, thus RC proposal is not to use it for Pb-Pb and pp ref. Will be discussed at coordination meeting. We provided one build with new CTF support included, to do validation tests, but all stable-sync builds we are doing now are without the new CTF coding improvements.
QC / Monitoring / InfoLogger updates:
- TPC has opened first PR for monitoring of cluster rejection in QC. Trending for TPC CTFs is work in progress. Ole will join from our side, and plan is to extend this to all detectors, and to include also trending for raw data sizes.
AliECS related topics:
- Extra env var field still not multi-line by default.
GPU ROCm / compiler topics:
- Found new HIP internal compiler error when compiling without optimization: -O0 make the compilation fail with unsupported LLVM intrinsic. Reported to AMD.
- Found a new miscompilation with -ffast-math enabled in looper folllowing, for now disabled -ffast-math.
- Must create new minimal reproducer for compile error when we enable LOG(...) functionality in the HIP code. Check whether this is a bug in our code or in ROCm. Lubos will work on this.
- Found another compiler problem with template treatment found by Ruben. Have a workaround for now. Need to create a minimal reproducer and file a bug report.
- Debugging the calibration, debug output triggered another internal compiler error in HIP compiler. No problem for now since it happened only with temporary debug code. But should still report it to AMD to fix it.
- New compiler regression in ROCm 5.6, need to create testcase and send to AMD.
TPC GPU Processing
- Random GPU crashes under investigation.
- Bug in TPC QC with MC embedding, TPC QC does not respect sourceID of MC labels, so confuses tracks of signal and of background events.
- Robert observed a segfault on the GPU in latest laser runs.
- Crash was with a custom-built O2 by TPC, not official release. Cannot reproduce the crash with official version.
- Investigating that, found a problem in the TPC timebin out of bounds check for triggered data, which is fixed in latest O2. In my test, this cut away plenty of clusters that were out of the trigger range. To be tested by TPC, and to be verified if the trigger handling is correct.
- PR with Cluster Error Parameterization extended with Streamers requested by Marian, and also cluster errors for IFC and when crossing CE.
- Online runs at low IR / low energy observe weird number of clusters per track statistics.
- Found a problem of incorrectly not disabling distortion maps when set to -1, thus similar observation in cosmics run where due to bogus applying 500 kHz distortion corrections.
- Still histograms do not look too good and checking if something is wrong, or if behavior in low IR data was always like this.
ANS Encoding
- New coding scheme finally merged in O2.
Issues currently lacking manpower, waiting for a volunteer:
- For debugging, it would be convenient to have a proper tool that (using FairMQ debug mode) can list all messages currently in the SHM segments, similarly to what I had hacked together for https://alice.its.cern.ch/jira/browse/O2-2108
- Redo / Improve the parameter range scan for tuning GPU parameters. In particular, on the AMD GPUs, since they seem to be affected much more by memory sizes, we have to use test time frames of the correct size, and we have to separate training and test data sets.