CMS Big Data Science Project
FNAL room: DIR/ Snake Pit-WH2NE - Wilson Hall 2nd fl North East
CERM room:
Instructions to create a light-weight CERN account to join the meeting via Vidyo:
If not possible, people can join the meeting by the phone, call-in numbers are here:
The meeting id is hidden below the Videoconference Rooms link, but here it is again:
- 10502145
Attendance: Alexeys, Cristina, Matteo, JimP, Ilia Cremer
News:
Oli and Saba not able to attend but Saba send status report of Nersc workflow.
-
Princeton Workflow:
-
Perform test next week and interface the scala code with histogrammar. Currently the code is able to skim and save objects and Cristina is dealing with getting information from external files (.json, .txt, .root, .C). Jim suggested to convert all of them to json to have a common format of the external files. Cristina will work this week on getting the code ready with Jim’s help.
-
Jim updated histogrammar including a remote viewer of the plots and bokah updates from Alexeys. Alexeys finished scala version for the main types of histograms: Histogram, Profile and Stack Plots. Just python documentation to be finished. Alexeys would like to test now any other operations or performance speedups needed, maybe including matrix operations.
-
There was some discussion on any heavy computational steps that the analysis may have that could use the implementation of new tools, e.g. multivariate analysis. Matteo clarified that the analysis is not using multivariate objects for the moment but would like to include them in the future (the validation is tricky since the training data (MC) does not represent the real data and that brings complications to validate the output of a BDT or TMVA by the collaboration)
-
Another step to consider could also be the calculation of weights. Matteo will clarify next week how much computational time does it consume and will try to follow Jim’s suggestion of doing the calculation on the fly while reproducing plots.
-
Some root files are corrupted. Cristina will check with Jim and transfer again those corrupted. Might thing on a future facility doing this automatically.
-
Matteo also suggested going back to the step of converting Bacon Production and will ask Nhan about the status.
-
-
Nersc workflow (summary of Saba’s report)
-
Cori is down for an upgrade for Phase II, so environment set up on Edison for now.
-
Couple of converted files are moved to NERSC ( 2 files are ~3GB and one is ~70MB)
-
Able to read the files into a Spark data frame (using Python API) and make simple plots. Can read in both layouts for the HDF5 files and can do both interactive and batch jobs.
-
H5Spark has some limitations as not being able to specify partitions per file while reading them.
-
Matteo will get in contact with Saba to update her about the analysis and hopefully she can take a look at the current scala code next week.
-
Saba plans to apply for an early project allocation at NERSC
-
Poster for the SC conference is due July 22nd. Saba will send out draft and request people to fill in.
-
-
Milestones for the next two weeks
-
Full scala test
-
Scala code conversion
-
Touchbase with Nhan
-
-
August 1st to August 5th: Alexeys will come to FNAL.