TY - JOUR
T1 - Cache topology aware computation mapping for multicores
AU - Kandemir, Mahmut
AU - Yemliha, Taylan
AU - Muralidhara, Sai Prashanth
AU - Srikantaiah, S.
AU - Irwin, Mary Jane
AU - Zhang, Yuanrui
PY - 2010/6
Y1 - 2010/6
N2 - The main contribution of this paper is a compiler based, cache topology aware code optimization scheme for emerging multicore systems. This scheme distributes the iterations of a loop to be executed in parallel across the cores of a target multicore machine and schedules the iterations assigned to each core. Our goal is to improve the utilization of the on-chip multi-layer cache hierarchy and to maximize overall application performance. We evaluate our cache topology aware approach using a set of twelve applications and three different commercial multicore machines. In addition, to study some of our experimental parameters in detail and to explore future multicore machines (with higher core counts and deeper on-chip cache hierarchies), we also conduct a simulation based study. The results collected from our experiments with three Intel multicore machines show that the proposed compiler-based approach is very effective in enhancing performance. In addition, our simulation results indicate that optimizing for the on-chip cache hierarchy will be even more important in future multicores with increasing numbers of cores and cache levels.
AB - The main contribution of this paper is a compiler based, cache topology aware code optimization scheme for emerging multicore systems. This scheme distributes the iterations of a loop to be executed in parallel across the cores of a target multicore machine and schedules the iterations assigned to each core. Our goal is to improve the utilization of the on-chip multi-layer cache hierarchy and to maximize overall application performance. We evaluate our cache topology aware approach using a set of twelve applications and three different commercial multicore machines. In addition, to study some of our experimental parameters in detail and to explore future multicore machines (with higher core counts and deeper on-chip cache hierarchies), we also conduct a simulation based study. The results collected from our experiments with three Intel multicore machines show that the proposed compiler-based approach is very effective in enhancing performance. In addition, our simulation results indicate that optimizing for the on-chip cache hierarchy will be even more important in future multicores with increasing numbers of cores and cache levels.
UR - http://www.scopus.com/inward/record.url?scp=77957584378&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=77957584378&partnerID=8YFLogxK
U2 - 10.1145/1809028.1806605
DO - 10.1145/1809028.1806605
M3 - Article
AN - SCOPUS:77957584378
SN - 1523-2867
VL - 45
SP - 74
EP - 85
JO - SIGPLAN Notices (ACM Special Interest Group on Programming Languages)
JF - SIGPLAN Notices (ACM Special Interest Group on Programming Languages)
IS - 6
ER -