2025년 7월에 PNNL의 ARTS 프로젝트에 합류했다. 합류한 뒤로 당사자들에게 들은 이야기와 그동안 읽은 논문, DOE 보고서들을 종합해서 ARTS가 어디서부터 이어져 온 계보인지 정리했다.

ARTS로 이어지는 계보의 뿌리는 둘이다. 하나는 MIT의 dataflow 연구자 Jack Dennis에서 시작한다. Dennis의 박사 제자 Guang Gao는[1] McGill 대학에서 dataflow와 폰노이만 모델을 결합한 멀티스레딩 아키텍처 EARTH를 만들었다[2]. 이후 델라웨어 대학으로 옮겨 CAPSL을 세웠고, 2011년 codelet 실행 모델을 제안했다[3]. CAPSL의 레퍼런스 구현체는 DARTS, Delaware Adaptive Run-Time System이다[4].

다른 하나는 DARPA에서 시작한다. DARPA는 2002년부터 HPCS(High Productivity Computing Systems) 프로그램으로 Cray, IBM, Sun에 각각 새 병렬 프로그래밍 언어 개발을 맡겼다. Cray는 Chapel, Sun은 Fortress, IBM은 X10을 냈다[5]. Chapel과 Fortress는 새 언어를 밑바닥부터 설계했고, X10은 Java를 순차 기반 언어로 쓰면서 UPC와 Titanium 같은 PGAS 언어에 비동기 태스크 병렬성을 결합했다. X10을 만든 IBM 왓슨 연구소의 Vivek Sarkar는 이후 Rice 대학으로 옮겨 X10을 Habanero 프로젝트로 확장했고[6], 2011년 Data-Driven Tasks 모델로 정리했다[7].

2012년 DOE ASCR의 X-Stack 프로그램은 아홉 개 프로젝트에 예산을 배정했다[8]. 그중 하나는 Intel의 Shekhar Borkar가 이끈 Traleika Glacier로, 공식 파트너는 Intel, Rice, UC San Diego, UIUC, Reservoir Labs, ET International, PNNL, University of Delaware였다[9]. 이 프로젝트에서 Rice의 Data-Driven Task 계보와 델라웨어의 codelet 계보가 EDT(Event-Driven Task), DataBlock, Event로 구성된 OCR, Open Community Runtime으로 정리됐다[10]. OCR 논문은 Related Work에서 이 두 계보를 모두 인용한다.

X-Stack의 또 다른 프로젝트는 DynAX였다. PI는 Guang Gao였고, 소속은 그가 세운 회사 ET International이었다[11]. PNNL과 델라웨어 대학도 참여했다. ETI는 codelet 모델을 상용화한 런타임 SWARM을 개발했다. 이 과정에서 관련 연구진들이 Traleika Glacier와 교류가 있었던 것으로 추정된다.

OCR은 여러 기관에서 구현됐다. Intel/Rice의 레퍼런스 구현과 비엔나 대학의 독자 구현(OCR-Vx)이 있고[13], PNNL의 구현은 P-OCR이다[12]. P-OCR의 핵심 저자인 Joshua Suetterlein과 Joshua Landwehr는 델라웨어 CAPSL에서 codelet 논문을 쓰던 연구자였고, PNNL로 옮겨 P-OCR을 만들었다. Guang Gao는 2020년까지 이 PNNL 논문들의 공저자로 남았다.

P-OCR을 만든 팀, Suetterlein과 Joseph Manzano, Andres Marquez가 2019년부터 ARTS를 만들었다[14]. ARTS는 Abstract Runtime System의 약자이며, P-OCR을 잇는 OCR 구현체로 EDT, DataBlock, Event 등 OCR 표준의 개념을 그대로 따른다. 2021년 LC-MEMENTO 논문은 Location Consistency(LC) 모델을 기반으로 GPU 지원을 추가했다[15].

2025년 3월 LLVM Performance Workshop에서 델라웨어 CAPSL과 PNNL이 함께 CARTS를 발표했다[16]. C는 컴파일러를 뜻하며, OpenMP 주석이 붙은 C/C++를 MLIR/LLVM을 거쳐 ARTS의 EDT와 DataBlock을 표현하는 dialect로 낮추는 컴파일러 프레임워크다.

현재는 내가 PNNL 동료들과 함께 ARTS와 CARTS를 연구하고 있으며, Rice 연구팀이 조지아텍으로 옮겨가 독자적인 방향으로 관련 연구를 계속하고 있는 것으로 안다[17].

전체 계보는 다음과 같다.

graph TD
    classDef delaware fill:#e8f0fe,stroke:#4a6fa5,color:#1a1a1a;
    classDef rice fill:#fff4e5,stroke:#c9860a,color:#1a1a1a;
    classDef intel fill:#f0f0f0,stroke:#555555,color:#1a1a1a;
    classDef pnnl fill:#e6f4ea,stroke:#2e7d32,color:#1a1a1a;
    classDef vienna fill:#f3e8fd,stroke:#7b3fa0,color:#1a1a1a;
    classDef here fill:#fdecea,stroke:#c0392b,color:#1a1a1a,stroke-width:3px;

    Dennis["Jack Dennis<br/>MIT dataflow"]:::delaware
    Gao["Guang Gao<br/>PhD 1986"]:::delaware
    Earth["EARTH / EARTH-MANNA<br/>McGill, 1990s"]:::delaware
    Codelet["Codelet execution model<br/>UDel CAPSL, 2011"]:::delaware
    Darts["DARTS<br/>UDel CAPSL, 2013"]:::delaware

    X10["X10<br/>IBM, 2005"]:::rice
    Habanero["Habanero project<br/>Rice, 2009"]:::rice
    DDT["Data-Driven Tasks/Futures<br/>Rice, 2011"]:::rice

    DynAX["DynAX (SWARM)<br/>ETI · PNNL · Reservoir · UIUC<br/>2012-2015"]:::delaware
    Traleika["Traleika Glacier (OCR)<br/>Intel · Rice · PNNL · UDel<br/>2012-2015"]:::intel

    XSOCR["XSOCR<br/>Intel/Rice"]:::intel
    POCR["P-OCR<br/>PNNL"]:::pnnl
    OCRVx["OCR-Vx<br/>Vienna"]:::vienna

    ARTS["ARTS<br/>PNNL, 2019-"]:::here
    CARTS["CARTS<br/>PNNL, 2025-"]:::here

    GT["Georgia Tech<br/>independent research"]:::rice

    Dennis --> Gao --> Earth --> Codelet
    Codelet --> Darts
    Codelet --> DynAX
    DynAX --> POCR

    X10 --> Habanero --> DDT

    Codelet --> Traleika
    DDT --> Traleika

    Traleika --> XSOCR
    Traleika --> POCR
    Traleika --> OCRVx

    POCR --> ARTS
    ARTS --> CARTS

    Habanero --> GT

참고문헌

[1] Wikipedia, “Guang Gao,” https://en.wikipedia.org/wiki/Guang_Gao
[2] H. Hum, O. Maquelin, K. Theobald, X. Tian, X. Tang, G. Gao et al., “A design study of the EARTH multiprocessor,” PACT 1995
[3] S. Zuckerman, J. D. Suetterlein, R. C. Knauerhase, G. R. Gao, “Using a ‘Codelet’ Program Execution Model for Exascale Machines: Position Paper,” EXADAPT ‘11, DOI: 10.1145/2000417.2000424
[4] J. Suettlerlein, S. Zuckerman, G. Gao, “An Implementation of the Codelet Model,” Euro-Par 2013, DOI: 10.1007/978-3-642-40047-6_63; J. D. Suetterlein, “DARTS: A runtime based on the Codelet execution model,” M.S. thesis, University of Delaware, 2014
[5] P. Charles et al. (incl. V. Sarkar), “X10: An Object-Oriented Approach to Non-Uniform Cluster Computing,” OOPSLA 2005; J. Dongarra, R. Graybill, W. Harrod et al., “DARPA’s HPCS Program: History, Models, Tools, Languages,” Advances in Computers 72, 2008, pp. 1-100
[6] R. Barik et al. (incl. Sarkar), “The Habanero Multicore Software Research Project,” OOPSLA Companion 2009
[7] S. Taşırlar, V. Sarkar, “Data-Driven Tasks and their Implementation,” ICPP 2011
[8] DOE ASCR, “ASCR X-Stack Portfolio,” https://science.osti.gov/ascr/Research/Computer-Science/ASCR-X-Stack-Portfolio
[9] S. Borkar (PI), “Traleika Glacier X-Stack, Final Scientific/Technical Report,” DOE OSTI 1226550, DOI: 10.2172/1226550, 2015
[10] T. G. Mattson et al. (17 authors), “The Open Community Runtime: A Runtime System for Extreme Scale Computing,” IEEE HPEC 2016, DOI: 10.1109/HPEC.2016.7761580
[11] G. Gao (PI), B. Meister, D. Padua, A. Marquez, “Final Project Report, DynAX,” DOE Award DE-SC0008716, 2015, https://www.osti.gov/servlets/purl/1238249
[12] J. Landwehr, J. D. Suetterlein, A. Marquez, J. Manzano, G. Gao, “Application characterization at scale: lessons learned from developing a distributed open community runtime system for high performance computing,” ACM Computing Frontiers 2016, DOI: 10.1145/2903150.2903166
[13] J. Dokulil, S. Benkner, “The OCR-Vx experience: lessons learned from designing and implementing a task-based runtime system,” Journal of Supercomputing 78, 2022, DOI: 10.1007/s11227-022-04355-0
[14] pnnl/ARTS GitHub repository, https://github.com/pnnl/ARTS
[15] K. Ranganath, J. S. Firoz, J. D. Suetterlein, J. B. Manzano Franco, A. Marquez, M. V. Raugas, D. Wong, “LC-MEMENTO: A Memory Model for Accelerated Architectures,” LCPC 2021, DOI: 10.1007/978-3-030-99372-6_5
[16] “CARTS: Enabling Event-Driven Task and Data Block Compilation for Distributed HPC,” 9th LLVM Performance Workshop @ CGO 2025, https://llvm.org/devmtg/2025-03/slides/carts.pdf
[17] Habanero Extreme Scale Software Research Laboratory, Georgia Tech, https://habanero.cc.gatech.edu/