Addis AI
Open Research

Open Research Project 01

Research groundwork underwayAmharic phase

Evaluating alignment and reliability in Amharic agricultural AI

This project measures how Amharic language adaptation and agricultural fine-tuning change model capability, refusal, uncertainty and locally relevant advice.

Underlying model-adaptation research is already underway. The funded evaluation phase begins after the minimum funding threshold or equivalent project-specific institutional funding is confirmed.

External funding toward this phase

$0

of $27,500 full target

Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.

$0Minimum useful scope $14,000
Minimum scope
$14,000
Full target
$27,500
Fund this research

Research question

How do continued pretraining, broad instruction tuning and agricultural specialization change the alignment and usefulness of an open model in Amharic?

Why this matters

Teams are adapting open models to languages and domains with limited evaluation of what changes at each stage. Agriculture makes the problem concrete because local conditions, missing context and confident errors can change the value of an answer.

What already exists

Public work this project builds on

Existing research

Existing adaptation work

Addis AI has previously carried out vocabulary extension, continued pretraining and supervised fine-tuning for Amharic. That work establishes the underlying adaptation pipeline and informs this study, but those earlier checkpoints are not treated as the controlled stage sequence for the research described here.

The new study will use a separately frozen experimental configuration so that changes introduced during language adaptation, instruction tuning and domain specialization can be evaluated consistently.

Research preview available on request.

The links below show Addis AI's broader open-source and benchmark track record. They are not presented as the controlled stage sequence for this study.

Stage comparison

One evaluation suite across four checkpoints

  1. 01

    Aligned instruction-tuned open checkpoint

    Start from an instruction-tuned open model with an existing safety and instruction-following baseline. The exact checkpoint and configuration will be selected and frozen in the public protocol before the main evaluation begins.

  2. 02

    Amharic language adaptation

    Apply the frozen Amharic language-adaptation configuration, then measure capability, factuality, language, local relevance, safety and alignment behavior again. The same agriculture suite is used so changes can be compared with the earlier checkpoints.

  3. 03

    Broad Amharic instruction tuning

    Teach general Amharic instruction following, then repeat the same evaluation suite. The same agriculture suite is used so changes can be compared with the earlier checkpoints.

  4. 04

    Domain specialization

    Apply agriculture or health supervised fine-tuning, then compare the change from every earlier stage. The same agriculture suite is used so changes can be compared with the earlier checkpoints.

What the evaluation protocol will freeze

The protocol will be published before the main stage-comparison evaluation begins.

  • Frozen research questions
  • Exact starting checkpoint and model configuration
  • Stage definitions and tokenizer policy
  • Training method and replay strategy at each stage
  • Chat template and decoding configuration
  • Scenario categories, evaluation rubrics and scoring definitions
  • Reviewer roles, judge policy and evaluation configuration
  • Statistical reporting approach
  • Separation of factuality, usefulness, safety and alignment outcomes

The protocol will state whether tokenizer changes are treated as part of the language-adaptation stage, separately controlled, or held fixed for the stage comparison.

Research questions

What the comparison will test

  • Which adaptation stage changes safety behavior most?
  • Does language adaptation weaken appropriate refusal or uncertainty?
  • Does domain specialization cause measurable alignment drift?
  • Are failures in Amharic missed by an equivalent English evaluation?
  • Can a small mitigation preserve both alignment and useful capability?
  • Does the model ask for missing local context before giving consequential agricultural advice?

Evaluation and metrics

Measures used at every stage

Capability and factuality

Whether the model gives a factually correct answer and retains useful task capability.

Language and local relevance

Whether it understands the Amharic input and produces an answer appropriate to the relevant local context.

Safety

Whether following the response could create meaningful harm.

Alignment behavior

Whether adaptation changes behaviors such as harmful compliance, inappropriate refusal, unsupported certainty, escalation or other safety-relevant behavior.

Project measures

  • Harmful compliance rate
  • Appropriate refusal rate
  • Escalation accuracy
  • Unsupported certainty and overconfidence
  • Domain factuality
  • Human usefulness
  • Amharic language quality
  • Change between model stages
  • Cross-language parity where applicable
  • Inter-rater agreement
  • Confidence intervals
  • Locally inappropriate recommendation rate
  • Missing-context sensitivity
  • Severity of harmful recommendations
  • Practical usefulness
  • Agronomist agreement

Scenario counts, reviewer redundancy, slice definitions and the statistical protocol will be fixed before the main evaluation. The public protocol will state exclusions and any later changes.

Reference standards

Authoritative reference hierarchy

  1. 1Ethiopian Ministry of Agriculture, ATI and relevant Ethiopian agricultural guidance
  2. 2Ethiopian agricultural research and locally applicable technical sources
  3. 3International references such as FAO where locally applicable
  4. 4Agronomist adjudication where sources are incomplete, contextual or conflicting

Reference selection and local-context decisions will be documented in the evaluation protocol.

Human and domain evaluation

Who reviews the model

Native Amharic evaluators test language quality and practical usefulness. Agronomists review factuality, missing context, severity and locally inappropriate recommendations.

What funding enables

  • Native Amharic evaluator time and agronomist review
  • Locally grounded scenario creation and preference labels
  • Contextual tests, red-team cases and repeated evaluation
  • Controlled mitigation experiments
  • Reproducibility work, documentation and public release

Execution plan

Who leads the work and when it runs

Co-leads
Biniyam Daniel
Kalkidan Demille
Research status
Research groundwork underway
Execution window
Approximately 8 weeks minimum
Approximately 10 to 12 weeks full scope
Next public artifact
Evaluation protocol
Funding trigger
Target start: within approximately 2 weeks after the minimum funding threshold is reached or equivalent project-specific institutional funding is confirmed.

Research team

  • 2 research engineers
  • Approximately 6 native Amharic evaluators
  • Agronomist reviewers
  • Rural participants where required by the full-scope evaluation design

Expected releases

Relative targets begin at project start. Actual dates will be added to the research log after funding is confirmed.

Evaluation protocol
Target within 4 weeks of project start and before the main stage-comparison evaluation begins
Minimum-scope report and evaluation release
Target within 8 weeks of project start
Full research report and final planned artifacts
Target within 12 weeks of project start

Mitigation experiments

Measure, identify, test, re-evaluate

A mitigation is tested only after the stage comparison identifies a failure or unwanted change. The affected metrics are then measured again.

  1. Adapt
  2. Measure
  3. Identify failures
  4. Test mitigation
  5. Re-evaluate
  6. Publish

Full scope

Base-distribution replay

Mix a controlled sample from the earlier training distribution into adaptation and test whether it limits drift.

Full scope

Safety data replay or injection

Add selected safety and alignment examples at the stage where evaluation finds a change, then re-run the suite.

Minimum and full scope

Domain-data quality filtering

Remove or reweight domain examples linked to factual, contextual or safety failures and compare the result.

Conditional

Adapter-capacity test

Test LoRA or a smaller adaptation capacity only if the first comparison suggests that parameter change is part of the failure.

Stretch

Preference alignment

Use a small DPO or preference-tuning experiment only if reviewer data and the remaining budget support it.

Research log

6 project checkpoints

Targets remain relative until funding is confirmed. When work starts, the log can add planned dates, actual completion dates, artifacts and delay or change notes without replacing the original target window.

01

Protocol and evaluation-set construction

Minimum scope

Protocol, source inventory, reviewer onboarding and evaluation-set construction.

Planned target
Weeks 1 to 2 after project start
Actual completion
Not completed
Artifact
Planned: Frozen protocol · Source inventory · Reviewer rubrics
Delay or change note
No change recorded
planned
02

Initial model-stage evaluation

Minimum scope

Baseline, CPT and broad-SFT evaluation.

Planned target
Weeks 3 to 4 after project start
Actual completion
Not completed
Artifact
Planned: Stage configurations · Initial comparison results
Delay or change note
No change recorded
planned
03

Domain and expert evaluation

Minimum scope

Agriculture-SFT evaluation, native review and agronomist review.

Planned target
Weeks 5 to 6 after project start
Actual completion
Not completed
Artifact
Planned: Native review · Agronomist review · Four-stage results
Delay or change note
No change recorded
planned
04

Minimum-scope analysis and release

Minimum scope

Failure analysis, initial results and minimum-scope public release.

Planned target
Weeks 7 to 8 after project start
Actual completion
Not completed
Artifact
Planned: Failure taxonomy · Initial report · Evaluation harness
Delay or change note
No change recorded
planned
05

Full-scope mitigation and context extension

Full-scope extension

Mitigation experiments and expanded contextual evaluation.

Planned target
Weeks 9 to 10 after project start
Actual completion
Not completed
Artifact
Planned: Mitigation configurations · Expanded contextual results
Delay or change note
No change recorded
planned
06

Repeat evaluation and full release

Full-scope extension

Repeat evaluation, analysis and full public release.

Planned target
Weeks 11 to 12 after project start
Actual completion
Not completed
Artifact
Planned: Repeat results · Full research report · Final planned artifacts
Delay or change note
No change recorded
planned

Funding scope

Minimum useful scope to full research target

The minimum funds a smaller but scientifically useful study with expert review. The expansion increases coverage, redundancy, mitigation testing, repeatability and release depth. Each work package appears once in the allocation below.

Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, field research, additional experiments and open release.

Minimum useful scope

$14,000

Full-scope expansion

$13,500

Full target

$27,500

Minimum useful scope

$14,000

A smaller but scientifically useful study with expert review, a four-stage comparison and a reproducible public release.

Evaluation protocol and Amharic evaluation-set construction
$3,000
Initial native-speaker evaluation
$2,500
Initial agronomist review
$2,500
Model-stage evaluation and failure analysis
$2,000
Training/evaluation compute and tooling
$1,500
Reproducibility, analysis and public release
$2,500
Total minimum scope
$14,000
Minimum share of full target51%
  • Frozen evaluation protocol
  • Initial Amharic agriculture evaluation suite
  • Baseline vs CPT vs broad-SFT vs agriculture-SFT comparison
  • Native-speaker review
  • Agronomist review
  • Initial failure taxonomy
  • Aggregate and per-slice results
  • Public methodology
  • Reproducible evaluation harness
  • Initial research report
  • Broad rural participant sampling
  • The full planned reviewer pool
  • Repeated independent evaluation rounds
  • The complete mitigation-ablation program
  • Extended field coordination
  • Additional languages

Full research target

$27,500

The additional funding increases evaluation coverage, reviewer redundancy, contextual evaluation, mitigation testing, repeatability, statistical confidence and reproducibility.

What the additional $13,500 adds

Expanded native and rural-context evaluation
$2,300
Expanded agronomist review and reviewer redundancy
$1,600
Field coordination and contextual evaluation
$2,400
Controlled mitigation experiments
$2,800
Repeat evaluation and statistical confidence work
$1,400
Additional engineering and reproducibility
$1,500
Documentation, release work and project contingency
$1,500
Total additional funding
$13,500
Expansion share of full target49%
Reconciliation: $14,000 minimum + $13,500 expansion = $27,500 full target.

Open deliverables

What we will publish

  • Evaluation protocol, failure taxonomy and scoring rubrics
  • Evaluation harness, scoring code and model-stage configurations
  • Aggregate and per-slice results with confidence intervals
  • Inter-rater agreement and cross-stage comparisons
  • Mitigation methods, results and reproduction instructions
  • Methodology, limitations and changes in scope
  • Datasets, checkpoints and source material where privacy, licensing and safety permit
  • We will publish null and negative results.

Beyond Addis AI

Why the result can be reused

The protocol and stage comparisons can be reused by teams running continued pretraining or supervised fine-tuning for other low-resource languages. There is little public evidence showing how alignment changes during this process.

Risks and limitations

Boundaries and planned responses

Project boundaries

  • The first phase evaluates Amharic only.
  • The work evaluates research systems. It does not replace agronomists, extension workers or locally applicable guidance.
  • Personal data and material that cannot be released lawfully or safely will not be published.

Risks and responses

Agricultural coverage is uneven
Report crop, task and context coverage. Do not generalize results beyond evaluated slices.
Formal Amharic dominates the evaluation set
Recruit reviewers across registers and report language coverage gaps.
A fluent answer hides a locally wrong recommendation
Score language separately from factuality, context and severity. Require agronomist review for consequential cases.
Open material creates licensing or safety problems
Track provenance and release only material that passes privacy, licensing and safety review.

Research updates

Project record

We publish progress, changes in scope, negative results and delays as they happen.

Milestone · Protocol and source review

Project scope published

The public scope, stage comparison, initial budget and planned outputs are available for review.

Findings: No project results are claimed at this stage.

Next step: Freeze the scenario, reviewer and statistical protocols before the main evaluation.

Funding

Support this project

Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships. Read the funding policy before contributing.

Addis AI funds the core engineering team and existing infrastructure. External funding primarily supports human evaluation, domain expertise, field research, additional experiments and open release.

Institutional funders can support a defined project work package through a separate project-specific agreement.

External funding toward this phase

$0

of $27,500 full target

Groundwork already funded and underway. Addis AI has funded the underlying model research and infrastructure; external funding expands this evaluation phase.

$0Minimum useful scope $14,000

Funding totals include only cleared contributions and confirmed project-specific grants or sponsorships.

Funding above the target

Funding above the target first expands Amharic sampling and independent re-review. Work in another language requires a separately published scope.

Choose a one-time amount

Payment is handled by our secure checkout provider. Addis AI does not receive or store your full card number.

Addis AI is a for-profit company. Contributions support the open research described on this page and are not represented as charitable or tax-deductible donations. No equity, investment return, profit sharing or tokens are offered.

By continuing, you acknowledge the Open Research funding policy.

Institutional and research contact

Grants, sponsorship, collaboration or research preview

contact@addisai.ch

This form is for research, grant and sponsorship conversations, not product demo requests.