mlx
MLX: An array framework for Apple silicon
파일 탐색기
최종 버전 다운로드 (.zip)- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- bug_report.md
- config.yml
- other.md
- build-sanitizer-tests.sh
- find-eligible-contributors.js
- setup+build-cpp-linux-fedora-container.sh
- update-bypass-list.js
- build_and_test.yml
- documentation.yml
- release.yml
- update_bypass_list.yml
- dependabot.yml
- pull_request_template.md
- autograd.cpp
- CMakeLists.txt
- compare_devices.cpp
- irregular_strides.cpp
- single_ops.cpp
- time_utils.h
- single_ops.py
- time_utils.py
- bench_gemm.py
- bench_gemv.py
- bench_mlx.py
- bench_torch.py
- compare.py
- README.md
- batch_matmul_bench.py
- block_masked_mm_bench.py
- compile_bench.py
- conv1d_bench.py
- conv2d_bench_cpu.py
- conv2d_train_bench_cpu.py
- conv2d_transpose_bench_cpu.py
- conv3d_bench.py
- conv3d_bench_cpu.py
- conv3d_train_bench_cpu.py
- conv3d_transpose_bench_cpu.py
- conv_bench.py
- conv_transpose_bench.py
- conv_unaligned_bench.py
- distributed_bench.py
- einsum_bench.py
- fft_bench.py
- gather_bench.py
- gather_mm_bench.py
- gather_qmm_bench.py
- hadamard_bench.py
- large_gemm_bench.py
- layer_norm_bench.py
- masked_scatter.py
- rms_norm_bench.py
- rope_bench.py
- scatter_bench.py
- sdpa_bench.py
- sdpa_vector_bench.py
- segmented_mm_bench.py
- single_ops.py
- slice_update_bench.py
- synchronize_bench.py
- time_utils.py
- extension.cmake
- FindCUDNN.cmake
- FindNCCL.cmake
- Findnvpl.cmake
- mlx_logo.pdf
- mlx_logo.png
- mlx_logo.svg
- mlx_logo_dark.pdf
- mlx_logo_dark.png
- mlx_logo_dark.svg
- m3-ultra-mesh-broken.png
- m3-ultra-mesh.png
- capture.png
- schema.png
- all-to-sharded-linear.png
- column-row-tp.png
- llama-transformer.png
- sharded-to-all-linear.png
- mlx_logo.png
- mlx_logo_dark.png
- module-base-class.rst
- nn-module-template.rst
- optimizers-template.rst
- ops.rst
- custom_metal_kernels.rst
- extensions.rst
- metal_debugger.rst
- metal_logging.rst
- mlx_in_cpp.rst
- data_parallelism.rst
- linear_regression.rst
- llama-inference.rst
- mlp.rst
- tensor_parallelism.rst
- distributed.rst
- functions.rst
- init.rst
- layers.rst
- losses.rst
- module.rst
- common_optimizers.rst
- optimizer.rst
- schedulers.rst
- array.rst
- cuda.rst
- data_types.rst
- devices_and_streams.rst
- distributed.rst
- export.rst
- fast.rst
- fft.rst
- linalg.rst
- memory_management.rst
- metal.rst
- nn.rst
- ops.rst
- optimizers.rst
- printoptions.rst
- random.rst
- transforms.rst
- tree_utils.rst
- compile.rst
- distributed.rst
- environment_variables.rst
- export.rst
- function_transforms.rst
- indexing.rst
- kv_cache.rst
- launching_distributed.rst
- lazy_evaluation.rst
- numpy.rst
- precision.rst
- quick_start.rst
- saving_and_loading.rst
- unified_memory.rst
- using_streams.rst
- conf.py
- index.rst
- install.rst
- .clang-format
- .gitignore
- .nojekyll
- Doxyfile
- index.html
- Makefile
- README.md
- requirements.txt
- CMakeLists.txt
- example.cpp
- README.md
- CMakeLists.txt
- distributed.cpp
- linear_regression.cpp
- logistic_regression.cpp
- metal_capture.cpp
- timer.h
- tutorial.cpp
- CMakeLists.txt
- eval_mlp.cpp
- eval_mlp.py
- README.md
- train_mlp.cpp
- train_mlp.py
- axpby.cpp
- axpby.h
- axpby.metal
- __init__.py
- bindings.cpp
- CMakeLists.txt
- pyproject.toml
- README.md
- requirements.txt
- setup.py
- test.py
- distributed_data_parallel.py
- distributed_tensor_parallel.py
- linear_regression.py
- logistic_regression.py
- qqmm.py
- .clang-format
- pocketfft.h
- binary.h
- broadcasting.cpp
- broadcasting.h
- buffer_cache.h
- CMakeLists.txt
- common.cpp
- compiled.cpp
- compiled.h
- copy.h
- hadamard.h
- load.cpp
- matmul.h
- metal_kernel.cpp
- metal_kernel.h
- quantized.h
- reduce.cpp
- reduce.h
- slicing.cpp
- slicing.h
- ternary.h
- unary.h
- utils.cpp
- utils.h
- bnns.cpp
- cblas.cpp
- simd_bf16.cpp
- simd_fp16.cpp
- simd_gemm.h
- accelerate_fp16_simd.h
- accelerate_simd.h
- base_simd.h
- math.h
- neon_fp16_simd.h
- simd.h
- type.h
- arange.h
- arg_reduce.cpp
- binary.cpp
- binary.h
- binary_ops.h
- binary_two.h
- cholesky.cpp
- CMakeLists.txt
- compiled.cpp
- compiled_preamble.h
- conv.cpp
- copy.cpp
- copy.h
- device_info.cpp
- device_info.h
- distributed.cpp
- eig.cpp
- eigh.cpp
- encoder.cpp
- encoder.h
- eval.cpp
- eval.h
- fft.cpp
- gemm.h
- hadamard.cpp
- indexing.cpp
- inverse.cpp
- jit_compiler.cpp
- jit_compiler.h
- lapack.h
- logsumexp.cpp
- luf.cpp
- make_compiled_preamble.ps1
- make_compiled_preamble.sh
- masked_mm.cpp
- matmul.cpp
- primitives.cpp
- qrf.cpp
- quantized.cpp
- reduce.cpp
- scan.cpp
- select.cpp
- slicing.h
- softmax.cpp
- sort.cpp
- svd.cpp
- ternary.h
- threefry.cpp
- threefry.h
- unary.cpp
- unary.h
- unary_ops.h
- add.cu
- arctan2.cu
- binary.cuh
- bitwise_binary.cu
- CMakeLists.txt
- divide.cu
- equal.cu
- greater.cu
- greater_equal.cu
- less.cu
- less_equal.cu
- log_add_exp.cu
- logical_and.cu
- logical_or.cu
- maximum.cu
- minimum.cu
- multiply.cu
- not_equal.cu
- power.cu
- remainder.cu
- subtract.cu
- conv.h
- gemm_conv.cu
- gemm_grouped_conv.cu
- copy.cuh
- copy_contiguous.cu
- copy_general.cu
- copy_general_dynamic.cu
- copy_general_input.cu
- atomic_ops.cuh
- binary_ops.cuh
- cast_op.cuh
- complex.cuh
- config.h
- cute_dequant.cuh
- fp16_math.cuh
- gather.cuh
- gather_axis.cuh
- gemm_sm70.cuh
- hadamard.cuh
- indexing.cuh
- qmm_naive.cuh
- qmm_sm80.cuh
- qmm_sm90.cuh
- scatter.cuh
- scatter_axis.cuh
- scatter_ops.cuh
- slice_update.cuh
- ternary_ops.cuh
- unary_ops.cuh
- utils.cuh
- block_mask.cu
- block_mask.h
- cublas_gemm.cpp
- cublas_gemm.h
- cublas_gemm_batched_12_0.cpp
- cublas_gemm_batched_12_9.cu
- gather_gemm.cu
- gather_gemm.h
- gemv.cu
- gemv.h
- grouped_gemm.h
- grouped_gemm_unaligned.cu
- fp_qmv.cu
- qmm.cu
- qmm.h
- qmm_naive.cu
- qmm_sm80.cu
- qmm_sm90.cu
- qmm_utils.h
- qmv.cu
- affine_quantize.cu
- convert_fp8.cu
- cublas_qqmm.cpp
- cublas_qqmm.h
- fp_quantize.cu
- fp_quantize.cuh
- mxfp8_quantize.cuh
- no_qqmm_impl.cpp
- nvfp4_quantize.cuh
- qqmm.cpp
- qqmm_impl.cpp
- qqmm_impl.h
- qqmm_utils.cu
- qqmm_utils.h
- quantized.cpp
- quantized.h
- quantized_utils.h
- all_reduce.cu
- col_reduce.cu
- init_reduce.cu
- reduce.cuh
- reduce_ops.cuh
- reduce_utils.cuh
- row_reduce.cu
- defines.cuh
- gemm.cuh
- mma.cuh
- tiles.cuh
- utils.cuh
- abs.cu
- arccos.cu
- arccosh.cu
- arcsin.cu
- arcsinh.cu
- arctan.cu
- arctanh.cu
- bitwise_invert.cu
- ceil.cu
- CMakeLists.txt
- conjugate.cu
- cos.cu
- cosh.cu
- erf.cu
- erf_inv.cu
- exp.cu
- expm1.cu
- floor.cu
- imag.cu
- log.cu
- log1p.cu
- logical_not.cu
- negative.cu
- real.cu
- round.cu
- sigmoid.cu
- sign.cu
- sin.cu
- sinh.cu
- sqrt.cu
- square.cu
- tan.cu
- tanh.cu
- unary.cuh
- allocator.cpp
- allocator.h
- arange.cu
- arg_reduce.cu
- bin2h.cmake
- binary_two.cu
- CMakeLists.txt
- compiled.cpp
- conv.cpp
- copy.cu
- cublas_utils.cpp
- cublas_utils.h
- cuda.h
- cuda_utils.h
- cudnn_utils.cpp
- cudnn_utils.h
- custom_kernel.cpp
- cutlass_utils.cuh
- delayload.cpp
- device.cpp
- device.h
- device_info.cpp
- dirs.cpp
- distributed.cu
- eval.cpp
- event.cu
- event.h
- fence.cpp
- fft.cu
- hadamard.cu
- indexing.cpp
- jit_module.cpp
- jit_module.h
- kernel_utils.cu
- kernel_utils.cuh
- layer_norm.cu
- load.cpp
- logsumexp.cu
- lru_cache.h
- matmul.cpp
- no_cuda.cpp
- primitives.cpp
- ptx.cuh
- random.cu
- reduce.cu
- rms_norm.cu
- rope.cu
- scaled_dot_product_attention.cpp
- scaled_dot_product_attention.cu
- scan.cu
- slicing.cpp
- softmax.cu
- sort.cu
- ternary.cu
- utils.cpp
- utils.h
- vector_types.cuh
- worker.cpp
- worker.h
- CMakeLists.txt
- copy.cpp
- copy.h
- device_info.h
- eval.h
- primitives.cpp
- scan.h
- slicing.cpp
- slicing.h
- includes.h
- indexing.h
- radix.h
- readwrite.h
- gather.h
- gather_axis.h
- gather_front.h
- indexing.h
- masked_scatter.h
- scatter.h
- scatter_axis.h
- ops.h
- reduce_all.h
- reduce_col.h
- reduce_init.h
- reduce_row.h
- steel_attention.h
- steel_attention.metal
- steel_attention_nax.h
- steel_attention_nax.metal
- attn.h
- loader.h
- mma.h
- nax.h
- params.h
- transforms.h
- steel_conv.h
- steel_conv.metal
- steel_conv_3d.h
- steel_conv_3d.metal
- steel_conv_general.h
- steel_conv_general.metal
- loader_channel_l.h
- loader_channel_n.h
- loader_general.h
- conv.h
- loader.h
- params.h
- steel_gemm_fused.h
- steel_gemm_fused.metal
- steel_gemm_fused_nax.h
- steel_gemm_fused_nax.metal
- steel_gemm_gather.h
- steel_gemm_gather.metal
- steel_gemm_gather_nax.h
- steel_gemm_gather_nax.metal
- steel_gemm_masked.h
- steel_gemm_masked.metal
- steel_gemm_segmented.h
- steel_gemm_segmented.metal
- steel_gemm_segmented_nax.h
- steel_gemm_segmented_nax.metal
- steel_gemm_splitk.h
- steel_gemm_splitk.metal
- steel_gemm_splitk_nax.h
- steel_gemm_splitk_nax.metal
- gemm.h
- gemm_nax.h
- loader.h
- mma.h
- nax.h
- params.h
- transforms.h
- integral_constant.h
- type_traits.h
- defines.h
- utils.h
- arange.h
- arange.metal
- arg_reduce.metal
- atomic.h
- bf16.h
- bf16_math.h
- binary.h
- binary.metal
- binary_ops.h
- binary_two.h
- binary_two.metal
- cexpf.h
- CMakeLists.txt
- complex.h
- conv.metal
- copy.h
- copy.metal
- defines.h
- dot.h
- dot.metal
- erf.h
- expm1f.h
- fence.metal
- fft.h
- fft.metal
- fp4.h
- fp8.h
- fp_quantized.h
- fp_quantized.metal
- fp_quantized_nax.h
- fp_quantized_nax.metal
- gemv.h
- gemv.metal
- gemv_masked.h
- gemv_masked.metal
- hadamard.h
- layer_norm.metal
- logging.h
- logsumexp.h
- logsumexp.metal
- quantized.h
- quantized.metal
- quantized_nax.h
- quantized_nax.metal
- quantized_utils.h
- random.metal
- reduce.h
- reduce.metal
- reduce_utils.h
- rms_norm.metal
- rope.metal
- scaled_dot_product_attention.metal
- scan.h
- scan.metal
- sdpa_vector.h
- searchsorted.h
- searchsorted.metal
- softmax.h
- softmax.metal
- sort.h
- sort.metal
- ternary.h
- ternary.metal
- ternary_ops.h
- unary.h
- unary.metal
- unary_ops.h
- utils.h
- allocator.cpp
- allocator.h
- binary.cpp
- binary.h
- CMakeLists.txt
- compiled.cpp
- conv.cpp
- copy.cpp
- custom_kernel.cpp
- device.cpp
- device.h
- device_info.cpp
- distributed.cpp
- eval.cpp
- event.cpp
- event.h
- fence.cpp
- fft.cpp
- hadamard.cpp
- indexing.cpp
- jit_kernels.cpp
- kernels.h
- logsumexp.cpp
- make_compiled_preamble.sh
- matmul.cpp
- matmul.h
- metal.cpp
- metal.h
- no_metal.cpp
- nojit_kernels.cpp
- normalization.cpp
- primitives.cpp
- quantized.cpp
- reduce.cpp
- reduce.h
- resident.cpp
- resident.h
- rope.cpp
- scaled_dot_product_attention.cpp
- scan.cpp
- slicing.cpp
- softmax.cpp
- sort.cpp
- ternary.cpp
- ternary.h
- unary.cpp
- unary.h
- utils.cpp
- utils.h
- CMakeLists.txt
- compiled.cpp
- device_info.cpp
- primitives.cpp
- allocator.cpp
- apple_memory.h
- CMakeLists.txt
- device_info.cpp
- eval.cpp
- event.cpp
- fence.cpp
- linux_memory.h
- primitives.cpp
- allreduce_bench.cpp
- CMakeLists.txt
- minimal_barrier.cpp
- minimal_cfg.cpp
- minimal_env.cpp
- group.h
- jaccl.cpp
- jaccl.h
- mesh.cpp
- mesh.h
- mesh_impl.h
- rdma.cpp
- rdma.h
- reduction_ops.h
- ring.cpp
- ring.h
- ring_impl.h
- tcp.cpp
- tcp.h
- threadpool.h
- types.h
- CMakeLists.txt
- README.md
- .gitignore
- CMakeLists.txt
- jaccl.cpp
- jaccl.h
- no_jaccl.cpp
- CMakeLists.txt
- mpi.cpp
- mpi.h
- mpi_declarations.h
- no_mpi.cpp
- CMakeLists.txt
- nccl.cpp
- nccl.h
- no_nccl.cpp
- CMakeLists.txt
- no_ring.cpp
- ring.cpp
- ring.h
- CMakeLists.txt
- distributed.cpp
- distributed.h
- distributed_impl.h
- ops.cpp
- ops.h
- primitives.cpp
- primitives.h
- reduction_ops.h
- utils.cpp
- utils.h
- CMakeLists.txt
- gguf.cpp
- gguf.h
- gguf_quants.cpp
- load.cpp
- load.h
- no_gguf.cpp
- no_safetensors.cpp
- safetensors.cpp
- bf16.h
- complex.h
- fp16.h
- half_types.h
- limits.h
- allocator.h
- api.h
- array.cpp
- array.h
- CMakeLists.txt
- compile.cpp
- compile.h
- compile_impl.h
- device.cpp
- device.h
- dtype.cpp
- dtype.h
- dtype_utils.cpp
- dtype_utils.h
- einsum.cpp
- einsum.h
- error.h
- event.h
- export.cpp
- export.h
- export_impl.h
- fast.cpp
- fast.h
- fast_primitives.h
- fence.h
- fft.cpp
- fft.h
- graph_utils.cpp
- graph_utils.h
- io.h
- linalg.cpp
- linalg.h
- memory.h
- mlx.h
- ops.cpp
- ops.h
- primitives.cpp
- primitives.h
- random.cpp
- random.h
- scheduler.cpp
- scheduler.h
- small_vector.h
- stream.cpp
- stream.h
- threadpool.h
- transforms.cpp
- transforms.h
- transforms_impl.h
- utils.cpp
- utils.h
- version.cpp
- version.h
- common.py
- config.py
- launch.py
- __init__.py
- activations.py
- base.py
- containers.py
- convolution.py
- convolution_transpose.py
- distributed.py
- dropout.py
- embedding.py
- linear.py
- normalization.py
- pooling.py
- positional_encoding.py
- quantized.py
- recurrent.py
- transformer.py
- upsample.py
- __init__.py
- init.py
- losses.py
- utils.py
- __init__.py
- optimizers.py
- schedulers.py
- __main__.py
- _reprlib_fix.py
- _stub_patterns.txt
- extension.py
- py.typed
- utils.py
- array.cpp
- buffer.h
- CMakeLists.txt
- constants.cpp
- convert.cpp
- convert.h
- cuda.cpp
- device.cpp
- distributed.cpp
- export.cpp
- fast.cpp
- fft.cpp
- indexing.cpp
- indexing.h
- linalg.cpp
- load.cpp
- load.h
- memory.cpp
- metal.cpp
- mlx.cpp
- mlx_func.cpp
- mlx_func.h
- ops.cpp
- print.cpp
- random.cpp
- random.h
- small_vector.h
- stream.cpp
- transforms.cpp
- trees.cpp
- trees.h
- utils.cpp
- utils.h
- __main__.py
- mlx_distributed_tests.py
- mlx_tests.py
- mpi_test_distributed.py
- nccl_test_distributed.py
- ring_test_distributed.py
- test_array.py
- test_autograd.py
- test_bf16.py
- test_blas.py
- test_compile.py
- test_constants.py
- test_conv.py
- test_conv_transpose.py
- test_device.py
- test_double.py
- test_einsum.py
- test_eval.py
- test_export_import.py
- test_fast.py
- test_fast_sdpa.py
- test_fft.py
- test_graph.py
- test_init.py
- test_linalg.py
- test_load.py
- test_losses.py
- test_memory.py
- test_nn.py
- test_ops.py
- test_optimizers.py
- test_quantized.py
- test_random.py
- test_reduce.py
- test_threads.py
- test_tree.py
- test_upsample.py
- test_vmap.py
- test_zero_copy.py
- allocator_tests.cpp
- arg_reduce_tests.cpp
- array_tests.cpp
- autograd_tests.cpp
- blas_tests.cpp
- CMakeLists.txt
- compile_tests.cpp
- creations_tests.cpp
- custom_vjp_tests.cpp
- device_tests.cpp
- einsum_tests.cpp
- eval_tests.cpp
- export_import_tests.cpp
- fft_tests.cpp
- gpu_tests.cpp
- linalg_tests.cpp
- load_tests.cpp
- ops_tests.cpp
- random_tests.cpp
- residency_tests.cpp
- scheduler_tests.cpp
- tests.cpp
- utils_tests.cpp
- vmap_tests.cpp
- .clang-format
- .gitignore
- .pre-commit-config.yaml
- ACKNOWLEDGMENTS.md
- AGENTS.md
- CITATION.cff
- CLAUDE.md
- CMakeLists.txt
- CODE_OF_CONDUCT.md
- CONTRIBUTING.md
- LICENSE
- MANIFEST.in
- mlx.pc.in
- pyproject.toml
- README.md
- setup.py
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/ml-explore/mlx
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd mlx
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. 공식 설치 스크립트
쉬움 추천사전 준비물
- Python 3 pip 명령어를 쓰려면 Python이 필요합니다.
pip install mlx
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mlx[cuda]
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mlx[cpu]
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
설치 후 새 터미널을 열고, 프로그램의 버전 확인 명령(예: --version)으로 정상 설치됐는지 확인하세요.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
3. CMake
보통사전 준비물
mkdir build && cd build
빌드 결과물을 담을 폴더를 만들고 그 안으로 이동합니다.
cmake ..
소스코드를 분석해 빌드 설정 파일을 생성합니다 (build 폴더 안에서 실행해야 함).
make
생성된 빌드 설정을 바탕으로 실제 컴파일을 진행해 실행 파일을 만듭니다.
build 폴더 안에 실행 파일이 생성됐는지 확인하고, 직접 실행해보세요 (예: ./build/앱이름).
4. Python
쉬움사전 준비물
pip install mlx
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mlx[cuda]
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
pip install mlx[cpu]
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
에러 메시지 없이 실행되고 터미널에 안내 문구가 출력되면 정상입니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
// repository documentation
Was this content helpful?
(0 ratings)
