letta-evals
Evaluation kit for testing stateful agents
파일 탐색기
최종 버전 다운로드 (.zip)- conventional-commits.yml
- deploy-leaderboard.yml
- e2e-tests.yml
- lint.yml
- publish.yml
- unit-tests.yml
- .release-please-manifest.json
- release-please-config.json
- evaluation-progress.png
- modal-sandbox.md
- custom_evaluators.py
- dataset.jsonl
- README.md
- suite.yaml
- task_1.py
- task_2.py
- task_3.py
- .gitignore
- custom_python_grader.py
- dataset.jsonl
- README.md
- setup.py
- suite.yaml
- dataset.jsonl
- README.md
- rubric.txt
- suite.yaml
- dataset.jsonl
- README.md
- rubric.txt
- suite.yaml
- dog.txt
- dataset.jsonl
- README.md
- rubric.txt
- suite.yaml
- create_agent.py
- dataset.jsonl
- inventory_tool.py
- README.md
- setup.py
- suite.yaml
- dataset.csv
- README.md
- rewards.py
- rubric.txt
- suite.logical-and.yaml
- suite.weighted-average.yaml
- core_memory_read.jsonl
- core-memory-read.yaml
- setup_agent.py
- core_memory_update.jsonl
- core-memory-update.yaml
- rubric.txt
- setup_agent.py
- filesystem_cloud.jsonl
- filesystem_code.jsonl
- addresses.txt
- bank_accounts.txt
- credit_cards.txt
- employments.txt
- insurance_policies.txt
- internet_accounts.txt
- medical_records.txt
- people.txt
- pets.txt
- vehicles.txt
- letta_file_bench.db
- agent_system_prompt.j2
- aggregation.md
- comparison_tiebreak.md
- cross_file_counting.md
- multi_entity_comparison.md
- multi_hop_chain.md
- negation.md
- question_quality_rubric.txt
- set_intersection.md
- temporal_reasoning.md
- __init__.py
- register_question_tool.py
- sql_execute_tool.py
- __init__.py
- config.yaml
- context.py
- display.py
- parallel.py
- question_generator.py
- README.md
- validate_questions.py
- create_code_dataset.py
- filesystem_cloud.yaml
- filesystem_code.yaml
- rubric.txt
- setup_cloud_agent.py
- leaderboard-chart.liquid
- anthropic.svg
- deepseek.svg
- google.svg
- minimax.svg
- mistralai.svg
- moonshotai.svg
- openai.svg
- z-ai.svg
- index.html
- letta-logo.svg
- styles.css
- .eleventy.js
- .gitignore
- package-lock.json
- package.json
- README.md
- dataset_select_use.csv
- dataset_use.csv
- __init__.py
- generate.py
- prompts.py
- utils.py
- .skills
- analyze_results.py
- create_dataset.py
- custom_extractor.py
- prompts.py
- rubric_skill_use.txt
- rubric_task_completion.txt
- suite_baseline.yaml
- suite_skill_select_use.yaml
- suite_skill_use.yaml
- generate_leaderboard_results.py
- leaderboard_filesystem_results.yaml
- leaderboard_skill_results.yaml
- README.md
- updates.md
- __init__.py
- hf.py
- loader.py
- __init__.py
- grading.py
- trace.py
- __init__.py
- builtin.py
- registry.py
- utils.py
- __init__.py
- base.py
- builtin.py
- prompt_utils.py
- rubric.py
- tool.py
- __init__.py
- results.py
- sample.py
- specs.py
- summaries.py
- __init__.py
- base.py
- dispatch.py
- Dockerfile
- modal.py
- __init__.py
- errors.py
- letta_code_target.py
- __init__.py
- base.py
- factory.py
- noop_progress.py
- progress_fields.py
- reducer.py
- rich_progress.py
- rich_renderer.py
- simple_progress.py
- state.py
- summary.py
- __init__.py
- cli.py
- constants.py
- decorators.py
- metrics.py
- pricing.py
- py.typed
- rewards.py
- runner.py
- streaming.py
- types.py
- utils.py
- __init__.py
- conftest.py
- test_cleanup.py
- test_cli_model_handle_override.py
- test_cli_single_sample.py
- test_cli_validate.py
- test_examples_e2e_live.py
- test_execution_trace_token_data.py
- test_extractors.py
- test_graders.py
- test_hf_dataset.py
- test_import_laziness.py
- test_letta_code_target_env_overrides.py
- test_letta_code_target_stream.py
- test_loader.py
- test_metrics.py
- test_modal_sandbox.py
- test_pricing.py
- test_rewards.py
- test_rich_progress.py
- test_rich_renderer.py
- test_rubric_grader.py
- test_runner_base_url.py
- test_runner_sandbox_dispatch.py
- test_runner_token_data.py
- test_streaming.py
- test_visualization_reducer.py
- test_visualization_summary.py
- .env.example
- .gitignore
- .gitmodules
- .pre-commit-config.yaml
- CHANGELOG.md
- CONTRIBUTING.md
- LICENSE
- pyproject.toml
- README.md
- uv.lock
// repository documentation
Was this content helpful?
(0 ratings)
