전체 글 139

6. Multicore Programming

[ MIT OpenCourseWare : Multicore Programming ] https://youtu.be/dx98pqJvZVk?si=2cdZid4RgwTfai_u Part 1. 왜 멀티코어 프로세서가 등장했는가1. Multicore processor의 기본 구조멀티코어 프로세서에서는 여러 개의 core가 같은 chip 위에 배치되어 있고, 이 core들은 shared memory에 접근할 수 있습니다.보통 각 core는 private cache를 가지고 있고, 여러 core가 공유하는 last-level cache도 있습니다. 예를 들어 L3 cache가 여기에 해당합니다. 또한 모든 core는 동일한 memory controller를 통해 메인 메모리에 접근하고, I/O에도 접근할 수 있습니..

2. Bentley Rules for Optimizing Work

[MIT OpenCourseWare] Bently Rules for Optimizing Work https://youtu.be/H-1-X9bkop8?si=XwVJWkN_mL2Pk0LG WorkDefinition.The work of a program (on a give input) is the sum total of all the operations executed by the program.프로그램의 work라는 것은 어떠한 입력이 주어졌을 때, 실행되는 모든 연산의 합. Optimizing Work알고리즘 설계는 문제 해결에 필요한 작업량을 드라마틱하게 줄일 수 있는데, 정렬 알고리즘의 O(n^2)와 O(n lg n)의 차이가 대표적인 예시이다. 하지만, 아래와 같은 컴퓨터 하드웨어의 복잡한 특성때..

[NDC25] MMO 서버에서 태스크 그래프를 활용한 확장성 있는 멀티스레드 아키텍처

[NDC25] MMO 서버에서 태스크 그래프를 활용한 확장성있는 멀티스레드 아키텍처 직접적으로 밝히지는 않았지만, DX의 MMO 서버 구조에 대한 멀티스레드 관리 방식을 간접적으로 볼 수 있음. https://youtu.be/6nYtK7kNiH8?si=OOdbc-UT5n5daXLn Contents1) Introduction- MMO 서버 도전 과제 2) Task Based Parallelism- Task를 이용한 작업 분배 문제 해결 3) Task Graph- Task간의 작업 흐름 관리 4) District ProcessorSeamless World를 위한 정적인 Task Graph 5) Exclusive Task- 공유 자원에 대한 배타적 접근을 보장하는 동적인 Task Graph 6) Limita..

The Cilk Runtime System

[MIT OpenCourseWare] The Cilk Runtime System https://youtu.be/Z7r4aAZ9Vqo?si=MdOHE_hw456cZ4nE The Cilk Runtime SystemP개의 프로세서에서, Tp - Cilk 스케줄러는 실행중인 프로그램을 프로세서 코어에 런타임에 동적으로 매핑시켜줌.- Cilk의 work-stealing 스케줄링 알고리즘은 효율적임이 입증됨. 핵심1) scheduling2) load balancing - Required Functionality- Performance Considerations- Implementating a Worker Deque- Spawning Computation- Stealing Computation- Synchron..

Parallel Storage Allocation

[MIT OpenCourseWare] Parallel Storage Allocation https://youtu.be/d5e_YJGXXFU?si=OjFtavN5i8PmsMXU Heap Storage in C지난 강의 Storage Allocation에서 다뤘던 부분 복습. Aligned allocationQ) aligned memory allocation을 쓰는 이유를 아시는분?A)1) 한 가지 이유는, 캐시라인에 정렬(align)되게 할당한다면, 2개의 캐시라인을 fetch해오지 않고 1개의 캐시라인만 fetch해올 수 있음. 따라서, 2번의 캐시미스대신 1번의 캐시미스만 발생.2) 또 다른 이유는 vectorization 연산이 2의 거듭제곱에 정렬(aligned)된 메모리 주소를 필요로 하기 때문..

Why should I learn Software Performance Engineering ?

[MIT OpenCourseWare] [1] 1. Introduction and Matrix Multiplication 일부you see once again the mention of "performance" in jobs is going up. and anecdotally, I can tell you, I hand one student who came to me after the spring, after he'd taken 6.172, and he said, "you know, I went and I applied for five jobs. and every job asked me, at every job interview, they asked me a question I couldn't have an..

Storage Allocation

[MIT OpenCourseWare] Storage Allocation https://youtu.be/nmMUUuXhk2A?si=1fDRy68DNBFxedky Stack Storage스택 메모리의 기본적인 할당/해제 방식 및 동작 구조 설명. Allocate x bytes할당량인 x만큼 sp(스택포인터)에 더한 뒤, 할당 전 시작 주소를 리턴.Free x bytes할당했던 x만큼 sp에서 빼면됨 대부분의 현대 컴파일러는 컴파일 타임에 stackoverflow를 검사하지 않는다. 왜냐하면 현대 컴퓨터에서 스택은 충분히 크기 때문.따라서 stackoverflow가 발생하더라도 segfault가 발생하고 후에 프로그램을 디버깅을 하면 된다. 따라서 효율성을 위해 stackoverflow를 확인하지 않음. St..