Manifund · cost–validity frontier
The Manifund page ranks proposals by thousands of pairwise LLM judgments. This companion eval asks what lexical features and frozen-embedding probes recover of that signal, per attribute — and how much the judges agree with each other in the first place. Forty proposals, one-line entities, rank statistics; directional, not decisive.
Open: rerun when the 1,269-pool phases land (pull_gold.py); add an absolute-scoring cheap-LLM rung; test whether the probe generalizes across pools.
Forty one-line proposal entities; 38/40 matched full descriptions. Rank correlations are descriptive at n=40: uncertainty intervals would be wide, so no p-values are presented. Description variants use the entity line for the two unmatched proposals except description-length correlations, which use the 38 matches only.
| Attribute | 31b–26b | 31b–Gemini | 26b–Gemini | Panel median |
|---|---|---|---|---|
| existential seriousness | 0.785 | 0.604 | 0.724 | 0.724 |
| technical depth | 0.882 | 0.732 | 0.789 | 0.789 |
| epistemic rigor of the underlying theory of change | 0.813 | 0.690 | 0.690 | 0.690 |
| counterfactual impact of a marginal dollar of funding | 0.702 | 0.453 | 0.452 | 0.453 |
| neglectedness of the problem the project addresses | 0.665 | 0.543 | 0.585 | 0.585 |
| tractability of the plan as proposed by this specific team | 0.712 | 0.568 | 0.629 | 0.629 |
| expected long-run impact on humanity's trajectory | 0.849 | 0.686 | 0.655 | 0.686 |
| scale of the problem the project addresses | 0.818 | 0.685 | 0.505 | 0.685 |
| empirical testability of its core claims | 0.806 | 0.799 | 0.836 | 0.806 |
| quality of quantitative reasoning in the proposal | 0.787 | 0.766 | 0.748 | 0.766 |
| legibility of the team's track record | 0.695 | 0.624 | 0.624 | 0.624 |
| feasibility given the stated team and resources | 0.634 | 0.751 | 0.596 | 0.634 |
| concreteness of proposed milestones | 0.779 | 0.741 | 0.753 | 0.753 |
| clarity of the theory of change | 0.642 | 0.698 | 0.516 | 0.642 |
| speed to a first observable result | 0.766 | 0.713 | 0.691 | 0.713 |
| room for more funding in this exact niche | 0.477 | 0.502 | 0.438 | 0.477 |
| cost-effectiveness per dollar spent | 0.490 | 0.467 | 0.534 | 0.490 |
| downside risk if the project succeeds at the wrong thing | 0.806 | 0.753 | 0.779 | 0.779 |
| probability of a net-negative outcome | 0.490 | 0.771 | 0.601 | 0.601 |
| robustness of the plan to its key assumptions being wrong | 0.349 | 0.541 | 0.208 | 0.349 |
| ambition | 0.671 | 0.752 | 0.579 | 0.671 |
| weirdness relative to mainstream research funding | 0.717 | 0.712 | 0.728 | 0.717 |
| interdisciplinarity | 0.701 | 0.759 | 0.639 | 0.701 |
| degree to which outputs are public goods | 0.706 | 0.589 | 0.681 | 0.681 |
| information value of simply running the project | 0.747 | 0.724 | 0.586 | 0.724 |
| replaceability by work others would do anyway | 0.571 | 0.484 | 0.469 | 0.484 |
| urgency of doing this now rather than in five years | 0.811 | 0.659 | 0.550 | 0.659 |
| maturity of the underlying field | 0.660 | 0.639 | 0.454 | 0.639 |
| sensitivity of the project's value to AI timelines | 0.632 | 0.745 | 0.513 | 0.632 |
| clarity of the writing itself | 0.664 | 0.720 | 0.483 | 0.664 |
| fundability by mainstream institutions | 0.602 | 0.573 | 0.373 | 0.573 |
| potential to become financially self-sustaining | 0.673 | 0.790 | 0.619 | 0.673 |
Pairwise medians: gemma31b–gemma26b 0.701; gemma31b–gemini 0.694; gemma26b–gemini 0.598. Overall panel-pair median 0.672. Seed-2 file contains 293 attributes; only 284 overlap the main run, not the stated 292. Finite, non-constant test–retest correlations: 280/284; median 0.464, IQR 0.349–0.559.
32-attribute gemma31b PCA: PC1 31.4%, PC2 17.6% of variance. PC1 sign is oriented so loadings sum positive.
| Attribute | PC1 loading |
|---|---|
| existential seriousness | 0.244 |
| technical depth | 0.136 |
| epistemic rigor of the underlying theory of change | 0.073 |
| counterfactual impact of a marginal dollar of funding | 0.131 |
| neglectedness of the problem the project addresses | 0.169 |
| tractability of the plan as proposed by this specific team | -0.249 |
| expected long-run impact on humanity's trajectory | 0.269 |
| scale of the problem the project addresses | 0.206 |
| empirical testability of its core claims | -0.057 |
| quality of quantitative reasoning in the proposal | 0.087 |
| legibility of the team's track record | -0.098 |
| feasibility given the stated team and resources | -0.229 |
| concreteness of proposed milestones | -0.035 |
| clarity of the theory of change | -0.126 |
| speed to a first observable result | -0.185 |
| room for more funding in this exact niche | 0.228 |
| cost-effectiveness per dollar spent | -0.086 |
| downside risk if the project succeeds at the wrong thing | 0.272 |
| probability of a net-negative outcome | 0.214 |
| robustness of the plan to its key assumptions being wrong | -0.230 |
| ambition | 0.259 |
| weirdness relative to mainstream research funding | 0.176 |
| interdisciplinarity | 0.127 |
| degree to which outputs are public goods | 0.002 |
| information value of simply running the project | 0.126 |
| replaceability by work others would do anyway | -0.185 |
| urgency of doing this now rather than in five years | 0.151 |
| maturity of the underlying field | -0.231 |
| sensitivity of the project's value to AI timelines | 0.258 |
| clarity of the writing itself | -0.156 |
| fundability by mainstream institutions | -0.006 |
| potential to become financially self-sustaining | 0.079 |
Subtle PCA: PC1 21.7%, PC2 11.4%. The source has 37,478/40,040 finite cells; missing cells were mean-imputed after per-attribute z-scoring, and 995/1,001 non-constant attributes entered PCA.
All probes use standardized embeddings, ridge alpha=1.0, and LOO predictions. The description variants use full descriptions for 38 entities and line-text fallback for two.
| Attribute | voyage4nano_line/gemma31b | voyage4nano_line/gemma26b | voyage4nano_line/gemini | voyage4nano_desc/gemma31b | voyage4nano_desc/gemma26b | voyage4nano_desc/gemini | gemma_line/gemma31b | gemma_line/gemma26b | gemma_line/gemini | gemma_desc/gemma31b | gemma_desc/gemma26b | gemma_desc/gemini |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| existential seriousness | 0.332 | -0.037 | 0.329 | 0.328 | -0.025 | 0.199 | 0.456 | -0.150 | 0.351 | 0.025 | -0.036 | 0.261 |
| technical depth | 0.477 | 0.551 | 0.545 | 0.486 | 0.494 | 0.546 | 0.527 | 0.709 | 0.669 | 0.486 | 0.587 | 0.608 |
| epistemic rigor of the underlying theory of change | 0.604 | 0.564 | 0.525 | 0.609 | 0.633 | 0.546 | 0.437 | 0.530 | 0.624 | 0.565 | 0.752 | 0.754 |
| counterfactual impact of a marginal dollar of funding | -0.002 | 0.143 | 0.034 | -0.134 | 0.311 | 0.211 | 0.214 | 0.279 | 0.250 | 0.045 | 0.303 | 0.198 |
| neglectedness of the problem the project addresses | 0.020 | 0.121 | 0.230 | -0.105 | 0.100 | -0.029 | -0.234 | -0.034 | -0.002 | -0.054 | 0.022 | -0.073 |
| tractability of the plan as proposed by this specific team | 0.287 | 0.118 | 0.335 | 0.010 | -0.052 | 0.384 | 0.203 | 0.239 | 0.370 | 0.016 | 0.092 | 0.274 |
| expected long-run impact on humanity's trajectory | 0.419 | 0.487 | 0.480 | 0.164 | 0.214 | 0.400 | 0.337 | 0.418 | 0.590 | 0.216 | 0.369 | 0.505 |
| scale of the problem the project addresses | 0.103 | 0.250 | 0.097 | 0.113 | 0.148 | 0.167 | 0.456 | 0.610 | 0.202 | 0.005 | 0.227 | 0.070 |
| empirical testability of its core claims | 0.369 | 0.334 | 0.446 | 0.409 | 0.453 | 0.577 | 0.452 | 0.635 | 0.749 | 0.507 | 0.436 | 0.704 |
| quality of quantitative reasoning in the proposal | 0.251 | 0.256 | 0.436 | 0.307 | 0.273 | 0.401 | 0.300 | 0.416 | 0.530 | 0.316 | 0.160 | 0.324 |
| legibility of the team's track record | -0.013 | -0.090 | -0.310 | 0.204 | -0.075 | -0.072 | 0.324 | 0.349 | -0.171 | 0.148 | -0.163 | -0.115 |
| feasibility given the stated team and resources | 0.426 | 0.146 | -0.011 | 0.445 | 0.031 | 0.179 | 0.339 | -0.027 | 0.054 | 0.505 | 0.198 | 0.089 |
| concreteness of proposed milestones | 0.090 | 0.172 | 0.087 | 0.094 | 0.242 | -0.011 | 0.446 | 0.506 | 0.456 | 0.172 | 0.259 | 0.071 |
| clarity of the theory of change | 0.452 | 0.272 | -0.026 | 0.022 | 0.083 | 0.032 | 0.539 | 0.319 | 0.299 | 0.294 | 0.134 | 0.420 |
| speed to a first observable result | 0.453 | 0.011 | -0.096 | 0.459 | 0.240 | 0.305 | 0.619 | 0.068 | 0.088 | 0.521 | 0.365 | -0.032 |
| room for more funding in this exact niche | 0.218 | 0.444 | 0.187 | -0.219 | 0.361 | 0.168 | -0.002 | 0.594 | 0.370 | 0.101 | 0.235 | 0.063 |
| cost-effectiveness per dollar spent | 0.129 | 0.056 | 0.166 | -0.022 | -0.052 | 0.431 | 0.072 | 0.346 | 0.286 | -0.070 | 0.275 | 0.497 |
| downside risk if the project succeeds at the wrong thing | 0.132 | 0.346 | 0.242 | -0.118 | 0.066 | 0.174 | 0.303 | 0.453 | 0.465 | -0.084 | 0.193 | 0.180 |
| probability of a net-negative outcome | 0.189 | 0.561 | 0.030 | -0.120 | -0.154 | -0.143 | 0.293 | 0.097 | -0.038 | -0.332 | -0.128 | -0.015 |
| robustness of the plan to its key assumptions being wrong | 0.145 | 0.338 | 0.414 | 0.306 | 0.204 | 0.307 | 0.264 | 0.063 | 0.385 | 0.155 | 0.005 | 0.395 |
| ambition | 0.347 | 0.075 | 0.180 | 0.270 | 0.259 | 0.435 | 0.262 | 0.129 | 0.046 | 0.105 | 0.250 | 0.124 |
| weirdness relative to mainstream research funding | 0.056 | 0.329 | 0.193 | 0.095 | 0.077 | 0.259 | 0.084 | 0.274 | 0.277 | 0.044 | -0.065 | 0.270 |
| interdisciplinarity | 0.122 | 0.067 | 0.507 | 0.221 | 0.244 | 0.530 | 0.184 | 0.141 | 0.534 | 0.249 | 0.228 | 0.189 |
| degree to which outputs are public goods | 0.329 | 0.605 | 0.259 | 0.247 | 0.279 | 0.063 | 0.450 | 0.491 | 0.343 | 0.148 | 0.315 | 0.105 |
| information value of simply running the project | 0.412 | 0.392 | 0.366 | 0.200 | 0.286 | 0.132 | 0.515 | 0.620 | 0.438 | 0.430 | 0.175 | 0.110 |
| replaceability by work others would do anyway | 0.222 | 0.191 | 0.005 | 0.530 | 0.190 | 0.113 | 0.227 | 0.339 | 0.030 | 0.504 | 0.368 | 0.164 |
| urgency of doing this now rather than in five years | 0.115 | 0.105 | 0.325 | 0.091 | -0.075 | 0.135 | 0.137 | 0.143 | 0.361 | 0.227 | -0.094 | 0.277 |
| maturity of the underlying field | 0.227 | 0.166 | 0.159 | 0.231 | -0.023 | 0.359 | 0.208 | 0.172 | 0.404 | 0.365 | 0.264 | 0.356 |
| sensitivity of the project's value to AI timelines | 0.339 | 0.116 | 0.423 | 0.251 | -0.078 | 0.529 | 0.149 | 0.294 | 0.556 | 0.303 | -0.089 | 0.718 |
| clarity of the writing itself | 0.233 | 0.297 | 0.005 | 0.284 | 0.258 | 0.100 | 0.033 | 0.152 | -0.158 | 0.285 | 0.077 | -0.080 |
| fundability by mainstream institutions | -0.020 | 0.105 | 0.104 | -0.037 | -0.017 | 0.256 | -0.212 | 0.071 | -0.126 | 0.197 | 0.052 | 0.067 |
| potential to become financially self-sustaining | 0.177 | 0.186 | 0.393 | 0.358 | 0.041 | 0.249 | 0.315 | 0.302 | 0.383 | 0.320 | 0.184 | 0.407 |
Finite correlations 995/1,001; Q1 0.026, median 0.199, Q3 0.348.
| Top attribute | rho |
|---|---|
| algorithm-shaped compromise | 0.712 |
| community-building instinct | 0.705 |
| fashionable-jargon half-life | 0.674 |
| bridge-building patience | 0.662 |
| mechanism before motivation | 0.656 |
| invention density | 0.655 |
| term-of-art precision | 0.648 |
| twitter-native compression | 0.645 |
| soft-power fluency | 0.641 |
| blame-preemption architecture | 0.638 |
| Bottom attribute | rho |
|---|---|
| accountability plumbing | -0.477 |
| upkeep budgeting | -0.477 |
| retraction grace | -0.559 |
| process maturity | -0.624 |
| urgency authenticity | -0.719 |
| sublimated ambition | -0.771 |
| audit-trail instinct | -0.865 |
| striving energy | -0.894 |
| structural inevitability | -0.949 |
| dialectical fairness | -0.975 |
Finite correlations 995/1,001; Q1 0.005, median 0.166, Q3 0.326.
| Top attribute | rho |
|---|---|
| newcomer welcome | 0.751 |
| beginner's-mind access | 0.706 |
| bedside manner | 0.697 |
| coalition breadth | 0.684 |
| data-availability instinct | 0.669 |
| term-of-art precision | 0.666 |
| curse-of-knowledge resistance | 0.663 |
| community-building instinct | 0.658 |
| soft-power fluency | 0.656 |
| field-notes register | 0.655 |
| Bottom attribute | rho |
|---|---|
| overdue-ness | -0.522 |
| vulnerability that isn't performance | -0.542 |
| urgency authenticity | -0.575 |
| example-to-principle ratio | -0.617 |
| victim-instrumentalization risk | -0.637 |
| sublimated ambition | -0.771 |
| dialectical fairness | -0.872 |
| audit-trail instinct | -0.883 |
| striving energy | -0.894 |
| structural inevitability | -0.949 |
Finite correlations 995/1,001; Q1 0.035, median 0.239, Q3 0.401.
| Top attribute | rho |
|---|---|
| invention density | 0.776 |
| falsifiable-promise density | 0.773 |
| blame-preemption architecture | 0.772 |
| beneficiary-voice presence | 0.749 |
| coalition breadth | 0.727 |
| twitter-native compression | 0.726 |
| mechanism before motivation | 0.721 |
| curse-of-knowledge resistance | 0.720 |
| guilt mobilization | 0.708 |
| community-building instinct | 0.699 |
| Bottom attribute | rho |
|---|---|
| chip-on-shoulder torque | -0.464 |
| polish without sterility | -0.509 |
| example-to-principle ratio | -0.560 |
| process maturity | -0.570 |
| evidence of killed darlings | -0.583 |
| striving energy | -0.671 |
| urgency authenticity | -0.766 |
| sublimated ambition | -0.829 |
| dialectical fairness | -0.872 |
| structural inevitability | -0.949 |
Finite correlations 995/1,001; Q1 -0.018, median 0.156, Q3 0.316.
| Top attribute | rho |
|---|---|
| discipline of its antifragility claims | 0.723 |
| blame-preemption architecture | 0.704 |
| term-of-art precision | 0.700 |
| field-notes register | 0.686 |
| first-principles freshness | 0.678 |
| twitter-native compression | 0.678 |
| reckless precision | 0.665 |
| stoic containment | 0.647 |
| curse-of-knowledge resistance | 0.643 |
| mechanism-design taste | 0.637 |
| Bottom attribute | rho |
|---|---|
| friction coefficient | -0.487 |
| upkeep budgeting | -0.492 |
| quotable-line frequency | -0.539 |
| deadline scent | -0.547 |
| playfulness under discipline | -0.616 |
| dialectical fairness | -0.667 |
| audit-trail instinct | -0.829 |
| sublimated ambition | -0.886 |
| striving energy | -0.894 |
| structural inevitability | -0.949 |
Single features are correlated directly with gemma31b gold. Description features use 38 matched records; the description ridge uses line fallback for two unmatched entities. Ridge uses alpha=1.0 and LOO.
| Attribute | Best single feature | |rho| | Line ridge | Description ridge |
|---|---|---|---|---|
| existential seriousness | line:flesch_reading_ease | 0.479 | 0.200 | -0.175 |
| technical depth | desc:capitalized_token_proportion | 0.532 | 0.311 | -0.366 |
| epistemic rigor of the underlying theory of change | line:word_count | 0.476 | 0.371 | -0.031 |
| counterfactual impact of a marginal dollar of funding | line:word_count | 0.417 | 0.056 | -0.286 |
| neglectedness of the problem the project addresses | line:word_count | 0.326 | -0.144 | -0.075 |
| tractability of the plan as proposed by this specific team | line:type_token_ratio | 0.314 | 0.013 | -0.061 |
| expected long-run impact on humanity's trajectory | line:flesch_reading_ease | 0.343 | 0.260 | -0.184 |
| scale of the problem the project addresses | line:number_percentage_density | 0.264 | -0.121 | -0.157 |
| empirical testability of its core claims | desc:capitalized_token_proportion | 0.507 | 0.247 | 0.156 |
| quality of quantitative reasoning in the proposal | line:word_count | 0.514 | 0.220 | 0.105 |
| legibility of the team's track record | line:word_count | 0.428 | 0.196 | -0.173 |
| feasibility given the stated team and resources | line:hedge_density | 0.257 | -0.101 | -0.382 |
| concreteness of proposed milestones | line:word_count | 0.564 | 0.479 | -0.256 |
| clarity of the theory of change | line:word_count | 0.510 | 0.275 | -0.482 |
| speed to a first observable result | line:capitalized_token_proportion | 0.331 | 0.053 | 0.058 |
| room for more funding in this exact niche | line:question_mark_count | 0.317 | 0.128 | 0.007 |
| cost-effectiveness per dollar spent | desc:number_percentage_density | 0.374 | 0.396 | 0.069 |
| downside risk if the project succeeds at the wrong thing | line:flesch_reading_ease | 0.317 | -0.023 | -0.140 |
| probability of a net-negative outcome | line:hedge_density | 0.271 | 0.282 | 0.130 |
| robustness of the plan to its key assumptions being wrong | desc:type_token_ratio | 0.289 | 0.179 | 0.064 |
| ambition | desc:question_mark_count | 0.377 | 0.255 | 0.113 |
| weirdness relative to mainstream research funding | desc:first_person_plural_ratio | 0.361 | -0.030 | 0.392 |
| interdisciplinarity | desc:first_person_plural_ratio | 0.405 | -0.142 | 0.243 |
| degree to which outputs are public goods | desc:capitalized_token_proportion | 0.328 | 0.002 | 0.095 |
| information value of simply running the project | line:word_count | 0.520 | 0.376 | -0.102 |
| replaceability by work others would do anyway | line:word_count | 0.363 | 0.010 | 0.053 |
| urgency of doing this now rather than in five years | line:word_count | 0.371 | 0.088 | -0.531 |
| maturity of the underlying field | desc:capitalized_token_proportion | 0.295 | -0.136 | 0.102 |
| sensitivity of the project's value to AI timelines | line:flesch_reading_ease | 0.340 | 0.020 | -0.071 |
| clarity of the writing itself | line:mean_word_length | 0.350 | 0.329 | -0.143 |
| fundability by mainstream institutions | desc:first_person_plural_ratio | 0.302 | 0.156 | 0.119 |
| potential to become financially self-sustaining | desc:first_person_singular_ratio | 0.460 | -0.272 | 0.330 |
Flags mark attributes where either available length measure has |rho(length, gold)| > 0.5.
| Attribute | Line↔gold | Desc↔gold | Line↔voyage4nano_line pred | Desc↔voyage4nano_line pred | Line↔voyage4nano_desc pred | Desc↔voyage4nano_desc pred | Line↔gemma_line pred | Desc↔gemma_line pred | Line↔gemma_desc pred | Desc↔gemma_desc pred | Line↔PC1 | Desc↔PC1 | Flag |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| existential seriousness | 0.142 | -0.117 | -0.091 | -0.360 | -0.215 | -0.212 | -0.165 | -0.242 | -0.221 | 0.017 | 0.127 | 0.030 | |
| technical depth | 0.426 | 0.204 | 0.228 | 0.214 | 0.230 | 0.259 | 0.206 | 0.129 | 0.324 | 0.234 | 0.127 | 0.030 | |
| epistemic rigor of the underlying theory of change | 0.476 | 0.337 | 0.312 | 0.313 | 0.345 | 0.406 | 0.175 | 0.185 | 0.514 | 0.342 | 0.127 | 0.030 | |
| counterfactual impact of a marginal dollar of funding | 0.417 | -0.015 | 0.032 | 0.134 | -0.073 | 0.072 | 0.202 | -0.054 | 0.038 | 0.199 | 0.127 | 0.030 | |
| neglectedness of the problem the project addresses | 0.326 | 0.218 | 0.016 | 0.103 | -0.168 | -0.031 | -0.036 | -0.116 | -0.053 | 0.083 | 0.127 | 0.030 | |
| tractability of the plan as proposed by this specific team | 0.256 | 0.218 | -0.077 | 0.086 | 0.109 | 0.191 | 0.409 | 0.021 | 0.132 | 0.151 | 0.127 | 0.030 | |
| expected long-run impact on humanity's trajectory | 0.263 | 0.096 | 0.163 | -0.168 | 0.211 | -0.066 | 0.128 | -0.175 | 0.102 | 0.057 | 0.127 | 0.030 | |
| scale of the problem the project addresses | -0.117 | -0.088 | 0.187 | -0.196 | -0.023 | -0.196 | -0.195 | -0.233 | -0.053 | -0.012 | 0.127 | 0.030 | |
| empirical testability of its core claims | 0.479 | 0.163 | 0.131 | 0.207 | 0.140 | 0.161 | 0.340 | 0.299 | 0.298 | 0.139 | 0.127 | 0.030 | |
| quality of quantitative reasoning in the proposal | 0.514 | 0.230 | 0.311 | 0.085 | 0.263 | 0.265 | 0.263 | 0.105 | 0.434 | 0.137 | 0.127 | 0.030 | LENGTH |
| legibility of the team's track record | 0.428 | 0.156 | 0.125 | 0.102 | 0.130 | -0.034 | 0.481 | 0.080 | 0.178 | 0.063 | 0.127 | 0.030 | |
| feasibility given the stated team and resources | 0.096 | -0.089 | -0.040 | 0.040 | 0.116 | -0.057 | 0.099 | 0.143 | 0.106 | -0.071 | 0.127 | 0.030 | |
| concreteness of proposed milestones | 0.564 | 0.103 | 0.123 | 0.055 | 0.067 | 0.103 | 0.440 | 0.071 | 0.254 | 0.061 | 0.127 | 0.030 | LENGTH |
| clarity of the theory of change | 0.510 | 0.022 | 0.269 | 0.190 | 0.085 | 0.137 | 0.576 | 0.198 | 0.254 | 0.294 | 0.127 | 0.030 | LENGTH |
| speed to a first observable result | 0.326 | 0.273 | 0.028 | 0.071 | 0.260 | 0.080 | 0.305 | 0.051 | 0.176 | 0.018 | 0.127 | 0.030 | |
| room for more funding in this exact niche | 0.040 | 0.034 | -0.092 | 0.064 | -0.056 | 0.101 | -0.159 | 0.114 | 0.035 | 0.139 | 0.127 | 0.030 | |
| cost-effectiveness per dollar spent | 0.332 | -0.193 | -0.044 | -0.024 | 0.055 | 0.071 | 0.290 | -0.002 | -0.124 | -0.050 | 0.127 | 0.030 | |
| downside risk if the project succeeds at the wrong thing | 0.157 | 0.034 | 0.001 | -0.308 | -0.091 | -0.180 | -0.096 | -0.223 | -0.136 | 0.073 | 0.127 | 0.030 | |
| probability of a net-negative outcome | 0.215 | -0.162 | 0.016 | -0.421 | -0.092 | -0.311 | 0.130 | -0.314 | -0.230 | -0.147 | 0.127 | 0.030 | |
| robustness of the plan to its key assumptions being wrong | 0.131 | 0.145 | -0.038 | 0.294 | 0.018 | 0.318 | 0.077 | 0.368 | 0.166 | 0.233 | 0.127 | 0.030 | |
| ambition | -0.018 | -0.141 | 0.181 | -0.204 | 0.144 | -0.227 | -0.028 | -0.216 | -0.040 | -0.025 | 0.127 | 0.030 | |
| weirdness relative to mainstream research funding | 0.208 | -0.013 | -0.129 | -0.203 | -0.131 | -0.192 | 0.037 | -0.219 | -0.233 | -0.155 | 0.127 | 0.030 | |
| interdisciplinarity | 0.121 | 0.162 | -0.127 | -0.186 | -0.238 | -0.181 | -0.310 | -0.278 | -0.207 | -0.030 | 0.127 | 0.030 | |
| degree to which outputs are public goods | 0.140 | 0.054 | 0.043 | -0.018 | 0.095 | -0.009 | -0.022 | -0.146 | 0.102 | 0.210 | 0.127 | 0.030 | |
| information value of simply running the project | 0.520 | 0.233 | 0.388 | 0.013 | 0.261 | 0.004 | 0.437 | -0.078 | 0.247 | 0.150 | 0.127 | 0.030 | LENGTH |
| replaceability by work others would do anyway | -0.363 | -0.247 | -0.380 | -0.266 | -0.308 | -0.467 | -0.260 | -0.015 | -0.380 | -0.539 | 0.127 | 0.030 | |
| urgency of doing this now rather than in five years | 0.371 | 0.174 | 0.089 | 0.068 | -0.083 | 0.185 | 0.045 | 0.256 | 0.186 | 0.419 | 0.127 | 0.030 | |
| maturity of the underlying field | -0.050 | -0.022 | -0.145 | -0.226 | -0.054 | -0.170 | -0.024 | -0.210 | -0.089 | -0.089 | 0.127 | 0.030 | |
| sensitivity of the project's value to AI timelines | 0.016 | 0.117 | 0.041 | -0.047 | 0.052 | -0.018 | -0.150 | -0.243 | 0.048 | 0.060 | 0.127 | 0.030 | |
| clarity of the writing itself | -0.078 | -0.126 | -0.067 | -0.028 | -0.298 | -0.003 | 0.118 | 0.041 | -0.138 | 0.188 | 0.127 | 0.030 | |
| fundability by mainstream institutions | -0.009 | -0.172 | 0.010 | 0.126 | -0.155 | -0.048 | 0.104 | 0.237 | 0.086 | 0.164 | 0.127 | 0.030 | |
| potential to become financially self-sustaining | 0.263 | -0.094 | 0.149 | 0.007 | 0.049 | -0.164 | 0.260 | -0.080 | -0.165 | -0.081 | 0.127 | 0.030 |
| Attribute | Panel median | Best lexical |rho| | Lexical ridge | Best embedding | Variant | Length rho |
|---|---|---|---|---|---|---|
| empirical testability of its core claims | 0.806 | 0.507 | 0.247 | 0.507 | gemma_desc | 0.479 |
| technical depth | 0.789 | 0.532 | 0.311 | 0.527 | gemma_line | 0.426 |
| downside risk if the project succeeds at the wrong thing | 0.779 | 0.317 | -0.023 | 0.303 | gemma_line | 0.157 |
| quality of quantitative reasoning in the proposal | 0.766 | 0.514 | 0.220 | 0.316 | gemma_desc | 0.514 |
| concreteness of proposed milestones | 0.753 | 0.564 | 0.479 | 0.446 | gemma_line | 0.564 |
| information value of simply running the project | 0.724 | 0.520 | 0.376 | 0.515 | gemma_line | 0.520 |
| existential seriousness | 0.724 | 0.479 | 0.200 | 0.456 | gemma_line | 0.142 |
| weirdness relative to mainstream research funding | 0.717 | 0.361 | -0.030 | 0.095 | voyage4nano_desc | 0.208 |
| speed to a first observable result | 0.713 | 0.331 | 0.053 | 0.619 | gemma_line | 0.326 |
| interdisciplinarity | 0.701 | 0.405 | -0.142 | 0.249 | gemma_desc | 0.121 |
| epistemic rigor of the underlying theory of change | 0.690 | 0.476 | 0.371 | 0.609 | voyage4nano_desc | 0.476 |
| expected long-run impact on humanity's trajectory | 0.686 | 0.343 | 0.260 | 0.419 | voyage4nano_line | 0.263 |
| scale of the problem the project addresses | 0.685 | 0.264 | -0.121 | 0.456 | gemma_line | -0.117 |
| degree to which outputs are public goods | 0.681 | 0.328 | 0.002 | 0.450 | gemma_line | 0.140 |
| potential to become financially self-sustaining | 0.673 | 0.460 | -0.272 | 0.358 | voyage4nano_desc | 0.263 |
| ambition | 0.671 | 0.377 | 0.255 | 0.347 | voyage4nano_line | -0.018 |
| clarity of the writing itself | 0.664 | 0.350 | 0.329 | 0.285 | gemma_desc | -0.078 |
| urgency of doing this now rather than in five years | 0.659 | 0.371 | 0.088 | 0.227 | gemma_desc | 0.371 |
| clarity of the theory of change | 0.642 | 0.510 | 0.275 | 0.539 | gemma_line | 0.510 |
| maturity of the underlying field | 0.639 | 0.295 | -0.136 | 0.365 | gemma_desc | -0.050 |
| feasibility given the stated team and resources | 0.634 | 0.257 | -0.101 | 0.505 | gemma_desc | 0.096 |
| sensitivity of the project's value to AI timelines | 0.632 | 0.340 | 0.020 | 0.339 | voyage4nano_line | 0.016 |
| tractability of the plan as proposed by this specific team | 0.629 | 0.314 | 0.013 | 0.287 | voyage4nano_line | 0.256 |
| legibility of the team's track record | 0.624 | 0.428 | 0.196 | 0.324 | gemma_line | 0.428 |
| probability of a net-negative outcome | 0.601 | 0.271 | 0.282 | 0.293 | gemma_line | 0.215 |
| neglectedness of the problem the project addresses | 0.585 | 0.326 | -0.144 | 0.020 | voyage4nano_line | 0.326 |
| fundability by mainstream institutions | 0.573 | 0.302 | 0.156 | 0.197 | gemma_desc | -0.009 |
| cost-effectiveness per dollar spent | 0.490 | 0.374 | 0.396 | 0.129 | voyage4nano_line | 0.332 |
| replaceability by work others would do anyway | 0.484 | 0.363 | 0.010 | 0.530 | voyage4nano_desc | -0.363 |
| room for more funding in this exact niche | 0.477 | 0.317 | 0.128 | 0.218 | voyage4nano_line | 0.040 |
| counterfactual impact of a marginal dollar of funding | 0.453 | 0.417 | 0.056 | 0.214 | gemma_line | 0.417 |
| robustness of the plan to its key assumptions being wrong | 0.349 | 0.289 | 0.179 | 0.306 | voyage4nano_desc | 0.131 |
The frontier uses the line-text lexical ridge; description-ridge results remain in section d as the extra-information variant.
Cost classes per item: LLM pairwise ≈ 12 comparisons × prompt tokens per entity per attribute; embedding = one local forward pass; lexical = free.
The main subtle run has 37,478 finite cells, 36 null scores, and 714 attributes with all 40 entities. Constant available scores make rank validity undefined for apology quality, corrosion resistance, self-doubt management, tu-quoque avoidance. Fewer than three finite entities occur for control of emphasis (n=2), inferential distance covered per paragraph (n=2), self-doubt management (n=2). The narrowest nonzero span is elephant-naming promptness (0.084, 4 unique values). Seed 2 has 293 attributes, with 9 absent from the main run and 284 overlapping attributes.