Integration guides · 2026-09-27

Testing Muse Spark Contributor in Turkish and Spanish applications

Build Turkish and Spanish checks for meaning, names, numbers and output format with Muse Spark Contributor. Evaluate each language without sending sensitive data.

A two-column LLMTR testing diagram brings Turkish and Spanish text into separate checks for meaning, names and numbers.

Set the data conditions before the example

Do not begin a Muse Spark Contributor evaluation with real customer conversations. Prompts and completions in this tier may be used by Meta for model training. For Turkish and Spanish tests, use material you wrote, have permission to share and that contains no personal information. A language test does not remove the data conditions.

This guide describes a test design for meta/muse-spark-1.2-contributor, not a live experiment or measured success rate. Standard tier and other versions are separate candidates. You can evaluate them on the same examples, but do not combine results without recording the model identifier. The question is not whether the model writes attractively in general, but whether your application behaves correctly in both languages.

Prepare inputs with the same meaning

Suppose your application summarizes a support message. Both language versions should contain the same event, date and requested action. Giving the Spanish example extra conditions while the Turkish example is shorter or clearer undermines the comparison. Have a person check semantic equivalence first. Do not count a condition altered during translation as a model error.

For a proposed test input, write a fictional delivery note: the item has not arrived and the user only wants information. The acceptance condition is that the summary must not invent delivery completion or a refund request. This is a suggested evaluation scenario, not an observed model response. Keep the input and acceptance condition in separate fields of the test record.

Score names, negation and numbers separately

One overall quality score can hide the defect that breaks your product. Record separate results for meaning, names, negation and numeric accuracy. Turkish characters and Spanish accents should be checked in both display and text processing. The model response is only part of the system: your application's handling of that text matters too.

Choose date and decimal formats before testing. The same digits can be misunderstood under different local conventions. Specify the intended presentation rather than asking the model to guess a country. If the source date is ambiguous, expect the ambiguity to remain. Adding certainty unsupported by the source is not successful localization, even when the sentence is grammatically polished.

Keep these checks separate in both languages
CheckExample of failure
MeaningTurning an information request into cancellation
NegationLosing the negative in not delivered
NamesReplacing a name with another word
Numbers and datesAdding a year to an ambiguous date
FormatChanging required fields or output language

Let the application validate its output contract

If your application expects an object or fixed fields, validate the format outside the response. Decide whether field names are translated, how missing values are represented and which types are accepted. Turkish and Spanish content may need to preserve the same technical keys. The requested format must match what your client actually accepts.

Do not delegate acceptance to a label the model adds to its own answer. If a required field is missing or has the wrong type, reject the result and define a correction path. If you intend to use strict schemas, verify support for the selected model and endpoint. Producing a JSON object and enforcing a particular schema are different capabilities.

One answer is not a language evaluation

Include ordinary cases, missing information and ambiguity in each language. Several attempts under the same evaluation conditions reveal variability. Report Turkish and Spanish separately rather than averaging away a serious defect in one language. Examine fields with costly errors independently of the overall score.

Record model identifier, date, settings, example revision and human assessment. Count outcomes without adding private material or customer answers to application logs. This guide contains no measured rate; report both numerator and denominator for your own trial. A small evaluation set does not represent every way a large user population will use the language, and a good average should not obscure a recurring critical error.

Base release decisions on failure types

An answer in the wrong language and an answer containing the wrong amount have different consequences. Decide in advance which failures block release. Reuse critical failure cases after corrections. Improving the score by adding only new, easier examples does not show that an earlier problem was fixed. Keep part of the evaluation set stable across changes.

Make human-review steps explicit in the first deployment. Retest both languages whenever the model or prompt changes. Adding a third language does not inherit validation from the first two. This method does not provide a ready-made quality guarantee; it makes accepted behavior and stopping conditions explicit. It also gives the next reviewer a clear reason for the decision rather than a collection of impressive-looking sample answers.

Frequently asked questions

Should I use customer messages in a Contributor test?

This guide recommends against it. Because of the training conditions, use material you may share that contains no personal or confidential data.

Is one successful Spanish response enough?

No. Include ordinary and ambiguous cases, repeated attempts and human review. Report results separately for each language.

Does JSON output prove schema compliance?

No. The application should validate fields and types. JSON generation and strict schema support are separate capabilities.

Related posts