Atlas:Analysis Challenge

Un article de lcgwiki.
Revision as of 17:21, 17 décembre 2008 by Chollet (talk | contribs) (First exercise on the FR Cloud (>= December 8th ))
Jump to: navigation, search

02/12/08 : E.Lançon, F.Chollet (Thanks to Cédric Serfon)

Information & Contact

Mailing list ATLAS-LCG-OP-L@in2p3.fr

Goals

  • measure "real" analysis job efficiency and turn around on several sites of a given cloud
  • measure data access performance
  • check load balancing between different users and different analysis tools (Ganga vs pAthena)
  • check load balancing between analysis and MC production

Required services @ T1

  • LFC catalog : lfc-prod.in2p3.fr
  • ATLAS Disk space : ATLASUSERDISK on T1 SE (fail-over for outputs in case of problems with T2 disk storage)

First exercise on the FR Cloud (December 2008)

Phase 1 : Site stress test oraganized by ATLAS and run centrally in a controlled manner (2 days)

DA challenges have been performed on IT and DE clouds in october 08. Proposition has been made to extend this cloud-by cloud challenge to the FR Cloud. See ATLAS coordination DA challenge meeting (Nov. 20)

First exercise will help to identify breaking points and bottlenecks. It is limited in time (a few days) and requires careful attention of site administrators during that period,in particular network (internal & external), disk, cpu monitoring. This first try (Stress tests) can be run centrally in a controlled manner. The testing framework is ganga-based.

  • Any site in the Tiers_of_ATLAS list can participate.
  • ATLAS coordination : Dan van der Ster and Johannes Elmsheuser
  • Details of Site Stress test : procedure, test conditions and targets
  • GlueCEPolicyMaxCPUTime >= 1440 (1 day , typical duration : 5 hours)
  • Jobs run under DN : /O=GermanGrid/OU=LMU/CN=Johannes_Elmsheuser
  • Results
  • Results available here : http://gangarobot.cern.ch/st/
  • Test 43 - Nov. 28
  • Scheduled Test 61 - December 8-10
Start Time: 2008-12-08 10:00:00
End Time: 2008-12-10 10:00:00
Test #61: http://gangarobot.cern.ch/st/test_61/
Sites: IN2P3-LPC_MCDISK GRIF-LPNHE_MCDISK TOKYO-LCG2_MCDISK IN2P3-CPPM_MCDISK 
Max Jobs Per Site: 300
Test #62: http://gangarobot.cern.ch/st/test_62/
Sites: TOKYO-LCG2_MCDISK
Max Jobs Per Site: 300
Output Dataset: user08.JohannesElmsheuser.ganga.sitetest.FR.081208.<sitename>
Input Type: DQ2_LOCAL
Input DS Patterns: mc08.*Wmunu*.recon.AOD.e*_s*_r5*tid* 
                   mc08.*Zprime_mumu*.recon.AOD.e*_s*_r5*tid* 
                   mc08.*Zmumu*.recon.AOD.e*_s*_r5*tid* 
                   mc08.*T1_McAtNlo*.recon.AOD.e*_s*_r5*tid* 
                   mc08.*H*zz4l*.recon.AOD.e*_s*_r5*tid* 
                   mc08.*.recon.AOD.e*_s*_r5*tid*

Tokyo will get 600 jobs over the 2 tests. Other sites will get 300 jobs, except CPPM which will get ~80 jobs because there are not many datasets available there.

Phase 2 : Pathena Analysis Challenge
  • Data Analysis exercice open to physicists with their favorite application
  • Physicists involved : Julien Donini, Arnaud Lucotte, Bertrand Brelier, Eric Lançon, LAL ?, LPNHE ?

Planning

  • Dec 8 : stop of MC production
  • Dec. 8-9: 1rst round with Tokyo, CPPM, LPC (LAN limited to 1Gbps), GRIF-LPNHE
  • Dec 17 : restart of MC production
  • Dec 14 : stop of MC production
  • Dec. 15-16 : 2nd round with LAPP, CC-IN2P3-T2 (to be confirmed), Tokyo, CPPM, LPC, possibly GRIF (SACLAY, IRFU, LPNHE), RO-07 and RO-02
  • Dec 17 : restart of MC production
  • Dec 17 : Beginning of Analysis Challenge (Phase 2)

Target and metrics

  • Nb of events : Few hundred up to 1000 jobs/site
  • Rate (evt/s) : up to 15 Hz
  • Efficiency (success/failure rate) : 80 %
  • CPU utilization : CPUtime / Walltime > 50 %