https://gitlab.com/Azuro721/trueperfect-ai (Link to Main Repo)

TruePerfect Sampler Settings

Turn any LLM into a perfected version of it only by using one-for-all universal sampler parameters.

This is my work of 1+ year; with trials and errors, I managed to achieve true perfection for any high-quality LLM.

Requirements:

Any Q6(K_M) model (Q5_K_M has issues), lower will not work correctly and there's no way to fix these issues, only temporarily by adjusting specific parameters.

CPU backend is required for exact replica of perfection for these settings!

Any other backend like VULKAN, cuBLAS, ROCm or anything besides CPU will cause EXTREMELY high level of overall degradation in any LLM and any combination of settings!

Smart Context should be disabled.

ContextShift should be disabled.

Flash Attention should be disalbed.

SWA is optional as can affect consistency stability.

Only KCPP 1.112.2 or below works correctly; any versions above has serious degradation in overall performance!

General and simplified info about specific sampler parameters:

General info + useful tips:
Temperature: Helps increasing creativity and usage of "smarter" words; has a strong dependence on Top-K (i.e. will produce very random and nonsensical results with incorrect combination of Top-K, TFS, Top-P and Repetition Penalty).
Top-K:

Helps expanding choices, diversity (variety), level of detail, logical structure of responses, stability over more complex user inputs, rules or overall character card and amount of simultaneous actions per turn (Top-K mainly does a good enhancement to correct combination of Repetition Penalty, Temperature, Top-P, TFS and Repetition Penalty Range); has a strong dependence on Temperature.

Will produce very chaotic, degrading, inconsistent, incoherent and nonsensical output with wrong value.

Repetition Penalty:

Helps repetitions on any models, controls creativity (more focused with lower values), consistency (better stability with lower values; needs very precise adjustment to achive "perfect" level of consistency comnined with other parameters like Top-P, Top-K, Temperature, Top-K and TFS), level of detail (with precise adjustment), complexity of developement, diversity and logic complexity (less subtle); very strict and needs extremely precise adjusting to achieve "perfect" level of consistency and overall performance.

Lower values (~1.07 and less) will produce more and more detailed output, with attention to things (smaller details), simultaneous events and more correct anatomical references for characters; tends to choose correct other details like species, body type and etc.

Higher values (~1.105 and more) will produce more focused outputs with descriptive way of phrasing and more directly focused outcomes based on events, choices and shifts; combine with higher Top-K and TFS to achieve maximum level of detail, length of output (per turn), attention, and slow down the overall progression.

Top-P: Helps to achieve correct outcomes based on events and complexity of actions; controls the way of phrasing, logic of actions and choices based on complexity, consistency, and overall logic; needs specific, highly fine-tuned values to achieve perfect coherence, smooth transitions, accurate descriptions, smooth flow, expected (less random) choices.
Top-A:

Helps to correct deviations related to Top-K.

Correct value will produce very creative, coherent and "rich" outputs (rich in diversity), with correct choices, consistency, actions and overall sense of logic; needs very precise adjustement.

Incorrect value will produce random, "rapidly-switching" and often nonsensical outputs, with complete mixup of details and sense of logic.

Has been removed since V3 due to inability to achieve "perfect" consistency.

TFS:

Helps to achieve perfect sense of logic with correct details and overall structural integrity of responses; will affect the way of phrasing, logic, consistency, and overall sense.

Needs extremely precise adjustement to achieve "perfect" results.

V3 is absolute upper limit for TFS.

Repetition Penalty Range:

Affects the overall complexity for in-depth developement of character and events; higher values (~70+) will help reducing any repetitions for Repetition Penalty (very effective against repetitions).

Higher values will also enhance the overall complexity and level of detail.

Incorrect values will have noticeable effects, such as very "rambly" (extremely long outputs without proper block division and logic without division); also the overall sense of logic.

Seed:

Fixed value helps to achieve much better consistency and flow.

Random will give more varied outputs with small chances of inconsistency, but also subtle reduction of any repetitions.

Strongly recommended to adjust Repetition Penalty Range and leave fixed.

Repetition Penalty Slope:

A "softer" way for Repetition Penalty; very unstable and inconsistent.

Values less than 1.0 only makes things incoherent and inconsistent.

Values higher than ~1.11 won't make any changes with strict settings, only slow down the generation.

Deactivated since V3 to achieve "perfect" consistency.

Min-P

Helps enhancing the stability and complexity of choices.

Deactivated since V3 due to negative effects in consistency and overall sense of logic.

Typical (Typicality):

More "intelligent" cutoff method; might output more "surprising" and varied outputs and "tails".

Used a fixed value of 1; any lower values has very negative impact on overall performance.

Presence Penalty:

Never had any positive results with it.

Affects the choices based on context, trained data and instructions only in a negative way; any value will cause issues ( even ones like 0.01 or lower).

Smoothing Factor:

Never had any positive results with it.

Increases the stability of outputs with very chaotic outcome; issue like nonsensical choices, actions, events and etc (even with low values like 0.002) are guaranteed.

Needs Smoothing Curve in order to work correctly.

Mirostat:

Never had any positive results with it.

Helps to steady the temperature (might replace) and dynamically adjusts the effective temperature based on more "surprising" tokens.

Tau helps to adjust the diversity; higher - more diverse and creative; lower - more deterministic.

Eta helps to increase the frequency of temperature updates; smaller - stable and slow; higher - faster but less coherent.

Only negative impact on consistency and overall performance with it.

Smoothing Curve:

Never had any positive results with it.

Dynamically adjusts the Penalty, temperature, probability to avoid sudden changes; 1 and lower.

Has stronger effect with higher Repetition Penalty.

Never had any positive results with it.

Values below 0.96 are not recommended.

Has no effect without Smoothing Factor.

Adaptive-P (KCPP's):

Avoid; Top-P only works with fixed value, and since it adjusts dynamically, only negative impact goes with it.

DRY:

Never had any positive results with it.

Helps to steady any repetitions, but never managed to get stability in consistency with longer progression.

Further progression will degrade logic, consistency and overall performance.

Use/Adjust Repetition Penalty Range instead.

XTC:

Direct enhancement to level of detail and attention to finer details, but never went well with logic and consistency.

Threshold - helps to cutoff the high-probability tokens (most likely), which mostly helps with lower Temperature.

Probability - the chance for Threshold to cutoff the desired most likely tokens.

Threshold should go lower than Probability in most cases for proper adjustment.

recommendations: do not use.

Ready-to-use AIO general settings:

Any AIO works well with FastForwarding (in most cases).

Works well (in most cases) if started from the second iteration (reloading same character card with same settings and sending the same input again; similar to 'Retry' right after the first word from the first output (not including the starting massage). Use only if FastForwarding is enabled.

Recommeded to preserve "consistency chain".

What is consistency chain and how to preserve it:

Consistency chain is consistent user inputs without any 'Retry' action or AI's output editing.

If you decide to edit AI's output, then you need to restore consistency chain.

To restore consistency chain, you need to send any letter instead of your usual input after finishing your editing, then stop right when AI starts writing anything, redo or remove whatever AI wrote, then send your desired input once again.

Main AIO settings + other useful info:
V1 **-CREATIVE-BALANCED-**

Very fine with very good creativity, level of detail (might skip some due to higher Top-A compared to V2), emotional connections, "surprising" outcomes and descriptions.

Optional Adjustments If Repetition Penalty 1.12082 outputs overly descriptive results, try improving the descriptions for character cards or altering the instructions, it will fix most of issues of such type.

In some cases, tends to confuse things like character names, species (mostly from trained data), pronouns, misuse (confuse) of User's actions, improper pronouns and etc (mostly due to low-probability token picks).

In some specific models (most likely), TFS 0.9551 might improve things even further, with more attention, level of creativity, and overall performance; do not use if you notice overly long descriptions (extremely long after each section).

V1 will lose coherence after ~11K tokens (up to ~14K tokens) or lower/higher depending on your character card, instructions and related things.

Proven to be less stable across higher-quality character cards and models.

V2 **-INSANE-DETAIL/ATTENTION-**

OVERKILL.

RECOMENDATIONS: use only if you want insane level of detail and attention; V1 is more preferred for general role play and good reactions, emotional connections, creativity and shorter outputs with well-preserved details. Use V2 only with very complex cases, or if you want extreme level of detail with very high attention and slow transisitons during certain events, as well as attention to very fine descriptions.

Maintains insane amount of details, attention, accuracy, and length: focused outputs, "surprising" outcomes and descriptions (noticeably (in some models) less compared to Top-A 0.07 but still generally good).

V2 will be more bland and less descriptive compared to V1 in favor of maximum stability and fixes related to general incoherence with character names (mostly from trained data) and improper pronouns (complete removal of very-low-probability token picks that produce rare misspellings and related issues).

Coherence will degrade after ~9K tokens.

Proven to be less stable across higher-quality character cards and models.

ASSISTANT MODE

TFS: 0.9551 Repetition Penalty: 1.02612

Insane for maximum accuracy for ASSISTANT-related tasks (personal assistant); Will be less creative in favor of attention.

Coherence will degrade after ~12K tokens.

If any issues occur (too detailed with incoherence) Decrease TFS to 0.8413 at cost of lower accuracy and level of detail, but noticeably better stability.
Optional Adjustments (might degrade stability and accuracy) Might provide very good results in specific models, but generally unstable.
Repetition Penalty 1.12082 Will improve creativity with slight reduction of accuracy.
Temperature 4.8 with Top-K 134 Will make outputs more lively, creative and "surprising", but might also increase chances of instabilities, such as overly-high description with sudden incoherence and overall degradation over higher amount of tokens.
Temperature 4.8 with Top-K 278 (284 / 296) Will make outputs more lively, creative, "surprising", descriptive and attentive, but tends to have higher chances of instability compared to Top-K134, with even faster degradation.
TFS 0.9551 (Special)

Will improve attention to details and overall performance in many aspects.

Special because tends to have more chances to work correctly across different models.

Lower Repetition Penalty to get even more insane attention and level of detail, but sacrifice a bit of creativity and "surprising" outcomes (especially with 1.02612).

Repetition Penalty 1.02612 is preferred as the lowest point; will output very attentive and detailed descriptions, with other things described earlier.

Use Top-K 134 for faster transitions and attention to more "surprising" moments (works (mostly) only with Top-A 0.07).

V3 **-BALANCED CLASSIC-**:

β˜…β˜…βœ­β˜†β˜† - Adaptation (ability to write excellently even with more complex character cards, rules and etc.)

β˜…β˜…β˜…β˜…β˜† - Stability (mostly logic stability across different occasions)

β˜…β˜…β˜…β˜…β˜… - Consistency stability (beyond 16K token limit)

β˜…β˜…β˜…β˜…β˜† - Writing style (mostly the way of phrasing and attention to overall more beautiful progression of events)

β˜…β˜…β˜†β˜†β˜† - Complexity (ability to engage and further develop character relationships, emotional states/connections, scene complexity and engagement between more characters; also expanded emotional palette under more complex events)

β˜…β˜…β˜…β˜†β˜† - Creativity

β˜…β˜…β˜…β˜†β˜† - Diversity

β˜…β˜…β˜…β˜…β˜… - Coherence

β˜…β˜…β˜†β˜†β˜† - Level of detail (Attention to finer details; but can go somewhat well in well-written character cards)

Very stable with excellent consistency under well-written character cards.

Will be higly repetitive if character card is poorly-written.

V3 **-BALANCED NARROW-**:

β˜…β˜…βœ­β˜†β˜† - Adaptation (Not adaptive in most average cases, but can go really well with good character card, user inputs and further progression of events)

β˜…β˜…β˜…β˜…βœ­ - Stability (Excellent due to narrow Repetition Penalty Range, assisted by Top-K134)

β˜…β˜…β˜…β˜…β˜… - Consistency stability (Excellent due to combination of narrower Repetition Penalty Range, Top-K134 and Repetition Penalty)

β˜…β˜…β˜…βœ­β˜† - Writing style (Goes OK in average; but very good with well-written character card)

β˜…β˜…β˜…β˜†β˜† - Complexity (But may go really well with well-written character cards and user inputs in combination with (in average) 2 characters)

β˜…β˜…β˜…β˜†β˜† - Creativity

β˜…β˜…βœ­β˜†β˜† - Diversity

β˜…β˜…β˜…β˜…β˜… - Coherence

β˜…β˜…β˜…β˜†β˜† - Level of detail (May go really well with combination of well-written character cards, user inputs and further progression of events)

Though stability is very high, the repetitions are as well with poorly-written character card.

This one needs well-written character card to work properly im most average cases.

V3_0 **-SLIGHTLY EXPANDED-**:

β˜…β˜…β˜…β˜†β˜† - Adaptation

β˜…β˜…β˜…β˜…β˜† - Stability

β˜…β˜…β˜…β˜…β˜… - Consistency stability

β˜…β˜…β˜…β˜…βœ­ - Writing style

β˜…β˜…βœ­β˜†β˜† - Complexity

β˜…β˜…β˜…βœ­β˜† - Creativity

β˜…β˜…β˜…βœ­β˜† - Diversity

β˜…β˜…β˜…β˜…β˜… - Coherence

β˜…β˜…βœ­β˜†β˜† - Level of detail

Slghtly better overall performance compared to V3.

Nearly no repetitions under any character card.

V3_1 **-EXCELLENT PERFORMANCE-**:

β˜…β˜…β˜…βœ­β˜† - Adaptation

β˜…β˜…β˜…β˜…βœ­ - Stability (May go unpredictable with very poorly-written character cards)

β˜…β˜…β˜…β˜…β˜… - Consistency stability

β˜…β˜…β˜…β˜…β˜… - Writing style (Second one with the most beautiful writing style)

β˜…β˜…β˜…β˜†β˜† - Complexity

β˜…β˜…β˜…βœ­β˜† - Creativity

β˜…β˜…β˜…β˜…β˜† - Diversity

β˜…β˜…β˜…β˜…β˜… - Coherence

β˜…β˜…β˜…β˜†β˜† - Level of detail

Good performance and the main recommended one across different character cards.

Mainly needs well-written character cards.

Though does not allow excellent emotional complexity over multiple characters.

V3_2 **-VERY GOOD ADAPTATION + WRITING STYLE-**:

β˜…β˜…β˜…β˜…βœ­ - Adaptation

β˜…β˜…β˜…β˜…βœ­ - Stability

β˜…β˜…β˜…β˜…β˜† - Consistency stability (Still holds perfect coherence over 16K+ tokens, but known to lose coherence over more complex user inputs)

β˜…β˜…β˜…β˜…β˜…β˜…- Writing style (Outstanding writing style across different occasions)

β˜…β˜…β˜…β˜…β˜† - Complexity

β˜…β˜…β˜…β˜…βœ­ - Creativity

β˜…β˜…β˜…β˜…β˜† - Diversity

β˜…β˜…β˜…β˜…βœ­ - Coherence (Can go slightly overboard in certain occasions)

β˜…β˜…β˜…β˜…β˜† - Level of detail

Best adaptation across different character cards, but Top-K0 causes far better diversity, sligtly better creativity and better attention to finer details and expands attention to advanced engagement between multiple characters (while Repetition Penalty is the main one who drives such things, but Top-K 0 enhances).

V3_3 **-EXCELLENT ADAPTATION-**:

β˜…β˜…β˜…β˜…β˜… - Adaptation

β˜…β˜…β˜…β˜…β˜… - Stability (Enhanced due to Top-K134, with Repetition Penalty being the main)

β˜…β˜…β˜…β˜…β˜… - Consistency stability (May go even better due to specific combination of Top-K and Repetition Penalty)

β˜…β˜…β˜…β˜…β˜…βœ­- Writing style (Excellent due to Repetition Penalty)

β˜…β˜…β˜…β˜…βœ­ - Complexity

β˜…β˜…β˜…β˜…β˜† - Creativity

β˜…β˜…β˜…β˜…βœ­ - Diversity

β˜…β˜…β˜…β˜…β˜… - Coherence (Excellent due to Top-K134)

β˜…β˜…β˜…β˜…β˜† - Level of detail

Enganced stability, but might seem shorter due to Top-K134.

V3_4 **-OUTSTANDING COMPLEXITY-**:

β˜…β˜…β˜…βœ­β˜† - Adaptation

β˜…β˜…β˜…β˜…β˜† - Stability (Goes unpredictable with poorly-written character cards, but extremely good with well-written)

β˜…β˜…β˜…β˜…β˜† - Consistency stability (Can lose consistency with poorly-written character cards, but handles complex user inputs relatively well)

β˜…β˜…β˜…β˜…β˜…βœ­- Writing style (Excellent due to Repetition Penalty and Repetition Penalty Range)

β˜…β˜…β˜…β˜…β˜…β˜…- Complexity (Outstanding complexity)

β˜…β˜…β˜…β˜…β˜† - Creativity

β˜…β˜…β˜…β˜…βœ­ - Diversity

β˜…β˜…β˜…β˜…β˜† - Coherence (Also goes unpredictable here due to poorly-written character cards)

β˜…β˜…β˜…β˜…βœ­ - Level of detail (Needs well-written character card for excellent results)

Insanely good complexity under well-written character cards, with outstanding engagament between characters, developement of events and emotional complexity between multiple characters and so on.

Requires well-written character card in order to work properly or at full potential.

To preserve the versatility, I would like to describe complete and specific sampler values below Additional fine-tuning:, to aim perfection for any case.

Here I will describe additional values for special cases, as well as dependencies across different sampler parameters (Like Temperature+Top-K)

Most of these fine-tuning methods are obsolete and will result in a much lower quality compared to any of V3.

Countless hours of testing proves that Temperature, Top-P, Repetition Penalty, Top-K, TFS, Repetition Penalty Range work together, and any distinct change will cause wide variety of issues, such as incoherense, loss of details, sense of logic and overall performance.

Some of the values for Repetition Penalty are noted and proven to be 'special' for stability, which might be considered for further fine-tuning alongside with others.

Additional fine-tuning (mostly obsolete):
Temperature+Top-K (only higher values):

Temperature is related to Top-K, and in order to achieve perfect part for this specific parameters, both of these needs to be adjusted.

For example Temperature 2.4 needs to have at least Top-K 134, increased by 72.

Acceptable values of Top-K for Temperature 2.4/4.8: 134, 206.

Different values will cause inconsistency, instability and other issues.

Acceptable values of Temperature for Top-K: 2.4, 4.8.

(1.2) is not prioritized mostly because in most cases it will output unsatisfactory results (e.g. bland and boring).

Lower temperature will be output less "exciting" and creative results (like less emotions, variety and predictability by one output), and might trigger repetitions, which can be mostly fixed by raising Top-K.

Higher Top-K will expand the attention to smaller details, and preserve attention to multiple simultaneous events, and also can fix smaller text-related issues (like with quotation marks, asterisks, hyphens and etc.)

Top-K 278, as described earlier, might cause overly descriptive results, which will most likely lead to incoherent results.

Top-K 206 is the more attentive one, which fits more with assistant tasks, as it will take away some of creativity, but tends to be more repetitive and might lead to incoherence.

Top-K 134 is the middle-balanced one, with better creativity, good level of detail and fine transitions. Recommeded one for in-character actions and strong roleplay scenarios NOT suitable for Top-A 0.0001.

Further experimentation with Top-K might not be possible, mostly due to logical limit for all settings combined.

Repetition Penalty:

Base value: 1.12082, which will output more creative, emotional, varied, smart and "exciting" results. But tends to have issues with asterisks and quotation marks; similar to 1.02612, but with more creativity, less descriptions, faster pace, but prone to issues if input has logical inconsistencies or lots of typos.

1.105 most stable across different combination of settings: specific value I found out during experimentation. Will output less "exciting" results, but fairly better compared to 1.05.

1.05 (not fine-tuned, not recommended): base value, which is widely used in various LLMs. Might output focused results with average creativity (better than 1.02612), but prone to be less stable compared to 1.105.

1.02612: very specific one the most stable one; will output very descriptive, attentive and expanded results. Will try to pay attention to noticeably more things compared to other variants. Will preserve character details and much more things as events go by. Great as an assistant.. Great for very complex instructions, very complex character cards and complex scenes. Great attention to multiple characters.

1.15 (not fine-tuned): tends to be more creative with shorter descriptions; might be better with Top-A 0.0001 and might be incoherent (might perform well on specific models).

1.23 (not fine-tuned): tends to be even more creative with slight shorter descriptions; might be better with Top-A 0.0001 and tends to have less chances to be coherent, but might perform well in rare cases (with specific models).

Other values (might output unstable results):

Feel free to experiment with these variants, and show any good results (if stable enough to be used for at least ~6K tokens).

1.02665: similar to 1.02695, but slightly more altered creativity, with interesting developement of events, more realistic responses from other characters; no issues so far; closest to the 1.02612 one; haven't been tested thoroughly.

1.02695: creative and more stable, with interesting developement of events, good "surprising" moments, realistic responses from other characters; no issues with initial tests; haven't been tested thoroughly.

1.0276/1.0277/1.0278: 1.0276 might provide repetitive results; 1.0277 is similar to 1.0285/1.0286, but with better stability and steady creativity; 1.0278 is more descriptive, but might be unstable with character details and fixed character type (like Pokemon).

1.0283 (decent): will be more direct and provide more realistic, violent scenes (if necessary), especially in uncensored models; good creativity, pays well attention to basic details, good progression of events, realistic responses from characters based on events, but noticeably shorter output and might be very "chatty".

1.0285/1.0286: will provide very interesting and creative responses, but will mess up some character details (mostly the ones that already in LLM's database) Some of them will output inconsistency with specific character parts, like body type, skin type and etc; also missing out certain details and skipping some important parts, be sure to include that and select the best one.

Might output quite stable results if used with Top-A 0.0001.

Top-P:

Base value: 0.915, which will output very attentive, consistent and stable results. Recommeded for all cases.

0.905: pays more attention to specific details, slightly less emotions, and very close to being repetitive.

0.95/0.97: very creative and unpredictable; might be used for better models, but generally less attentive (might perform well on higher-quality models (12B+)).

Top-A:

Base value: 0.07, which will output very consistent, generally stable results, with smooth transitions and relatively good attention to most details. Recommeded for all cases.

0.0001: will output insane amount of details, attention, accuracy and other things described in V2.

Other values (experimental):

0.043725: more attention to anatomy, but unstable and tends to be unpredictable. Works better with Repetition Penalty 1.02612 or V2 but degrades.

0.2025: more creative, descriptive, "exciting" and emotional, but tends to skip some details. Less accurate and have lots of issues with high Temperature; not suitable to be used generally, only to get initial creative inputs.

Repetition Range:

Base value: 64, which is the least one that will output overall better, consistant and descriptive results. Recommeded for all cases.

128: will output more descriptive results, but tends to be repetitive. Not recommended, but cab be used to adjust initial responses.

Seed:

Use fixed seed to improve consistency alot. Feel free to use these fixed values: 253991 main one

372205 - second one

309090 - third one

680079

637001

608575

132458

Repetition Slope:

Base value: 1.12, which will help to max out most things like consistency, level of detail and etc. Recommeded for all cases. Higher values won't affect outputs at all, only might cause issues with slowdowns.

All values below 1 are unstable and will cause very random issues.

TFS:

Base value: 0.8413, which will output very smart, attentive and smooth outputs. Recommeded for all cases.

0.9551: will output extremely descriptive and attentive results; the best one as an assistant; might cause issues in some LLMs as described earlier.

Higher TFS will lessen repetitions for lower Top-P, but V3 has absolute upper limit for TFS.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support