SentenceTransformer

This is a sentence-transformers model trained. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Maximum Sequence Length: 192 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'query: aayakar praadhikaran dvaaraa apiil karane ke lie kaun sii kar chhuut upalabdh hai?',
    'passage: धारा 373: आयकर प्राधिकरण द्वारा अपील दाखिल करना 373. (1) बोर्ड इस अध्याय के प्रावधानों के तहत किसी आयकर प्राधिकरण द्वारा अपील दाखिल करने के उद्देश्य से, इस अध्याय के प्रावधानों के तहत किसी भी आयकर प्राधिकरण द्वारा अपील दाखिल करने को विनियमित करने के उद्देश्य से, समय-समय पर अन्य आयकर प्राधिकरणों को आदेश, निर्देश या निर्देश जारी कर सकता है, यह ऐसे मौद्रिक सीमाओं को निर्धारित करने से नहीं रोक सकता है, जो वह किसी भी आयकर वर्ष के लिए किसी भी आयकर प्राधिकरण के मामले में किसी भी मुद्दे पर अपील नहीं कर सकता है। (2) जहां उप-धारा (1) के तहत जारी आदेशों, निर्देशों या निर्देशों का पालन करते हुए आयकर प्राधिकरण ने किसी भी मामले में किसी भी मुद्दे पर अपील नहीं दायर की है, तो यह इस तरह के किसी भी प्राधिकरण को किसी अन्य आयकर वर्ष के लिए किसी भी मामले में याचिका दायर करने से नहीं रोक सकता है। (3) जहां उप-धारा (1) के तहत किसी आयकर प्राधिकरण द्वारा या निर्देशों के तहत किसी भी मामले में याचिका दायर नहीं की गई है, और जहां किसी भी मामले में अपील करने वाले पक्ष द्वारा या याचिका दायर नहीं की गई है, उस मामले में अदालत द्वारा किसी भी याचिका दायर नहीं की जाएगी। (1)',
    'passage: Section 367: Appeal to Supreme Court\n\n367. An appeal shall lie to the Supreme Court from any judgment of the High \n Court delivered on an appeal made to High Court in respect of an order \npassed under section 363 in any case which the High Court certifies to be fit for \nappeal to the Supreme Court.\nHearing before Supreme Court.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.2213, 0.1421],
#         [0.2213, 1.0000, 0.3914],
#         [0.1421, 0.3914, 1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 4,000 training samples
  • Columns: anchor, positive, and negative
  • Approximate statistics based on the first 100 samples:
    anchor positive negative
    type string string string
    modality text text text
    details
    • min: 11 tokens
    • mean: 31.81 tokens
    • max: 104 tokens
    • min: 29 tokens
    • mean: 164.93 tokens
    • max: 192 tokens
    • min: 4 tokens
    • mean: 132.9 tokens
    • max: 192 tokens
  • Samples:
    anchor positive negative
    query: तलाशी और जब्त के लिए किस प्रकार की कर छूट उपलब्ध है? passage: Section 247: Search and seizure

    247. (1) Where the competent authority, in consequence of information in his

    possession, has reason to believe that—

    ( a) any person to whom a summons under section 131(1) or a notice under

    section 142(1) of the Income-tax Act, 1961 (43 of 1961) or summons

    under section 246(1) or a notice under section 268(1) of this Act,––

    ( I) was issued to produce, or cause to be produced, any books of account

    or other documents, or any information in electronic form or on

    a computer system, has omitted or failed to produce, or cause to

    be produced, such books of account or other documents or such

    information as required by such summons or notice; or

    ( II) has been issued or might be issued, will not, or would not, produce

    or cause to be produced, any books of account or other documents,

    or any information in electronic form or on a computer system

    which will be useful for, or relevant to, any proceedings under the

    Income-tax Act, 1961 (43 of 1...
    passage: धारा 296: धारा 63[(1) धारा 286 के प्रावधानों के बावजूद, धारा 294 के तहत आदेश उस तिमाही के अंत से अठारह महीने के भीतर पारित किया जाएगा जिसमें तलाशी शुरू की गई थी या तलाशी ली गई थी.] (2) जहां तलाशी शुरू की गई थी या तलाशी ली गई थी, और संबंधित खंड अवधि की कुल अज्ञात आय के आकलन या पुनर्मूल्यांकन के दौरान धारा 166 (1) के तहत कोई संदर्भ दिया गया है, ऐसे मूल्यांकन या पुनर्मूल्यांकन प्रक्रिया को पूरा करने के लिए उपलब्ध अवधि को बारह महीने तक बढ़ाया जाएगा। (3) उपधारा (1) के तहत समय सीमा की गणना में, उस समय अवधि (एक सौ और अस्सी दिनों से अधिक नहीं) जो उस तारीख से शुरू होती है जिस दिन ऐसी तलाशी शुरू की गई है या एक पुनर्मूल्यांकन किया गया है और जहां धारा 166 (1) के तहत प्रदान की गई अवधि की समाप्ति के दौरान धनराशि का पुनर्मूल्यांकन किया गया है, और जब धारा 261 के अंत के अंत के लिए ऐसी अवधि समाप्त हो जाती है, तो ऐसे अनुच्छेद (2) के तहत किसी अन्य व्यक्ति द्वारा इस तरह के मूल्यांकन की अवधि के अंत तक या पुनर्मूल्यांकन की अवधि के लिए संदर्भित किया जाएगा, यदि धारा 29 (5) के अंत तक किसी अन्य व्यक्ति ...
    query: aay kii kam riporting aur galat riporting ke lie jurmaanaa ke lie kaun sii kar chhuut upalabdh hai? passage: Section 439: Penalty for under-reporting and misreporting of income

    (12) The tax payable in respect of the under-reported income shall be—

    ( a) where no return of income has been furnished or where return has been

    furnished for the first time under section 280 and the income has been

    assessed for the first time, the amount of tax calculated on the under-re-

    ported income as increased by the maximum amount not chargeable to

    tax as if it were the total income;

    ( b) where the total income determined under section 270(1)(a) or assessed,

    reassessed or recomputed in a preceding order is a loss, the amount of

    tax calculated on the under-reported income as if it were the total income;

    ( c) in any other case, determined as follows—

    (X – Y)

    where,—

    X = the amount of tax calculated on the under-reported income

    as increased by the total income determined under section

    270(1)(a) or total income assessed, reassessed or recomputed

    in a preceding order as if it were the total ...
    passage: Section 367: Appeal to Supreme Court

    367. An appeal shall lie to the Supreme Court from any judgment of the High
    Court delivered on an appeal made to High Court in respect of an order
    passed under section 363 in any case which the High Court certifies to be fit for
    appeal to the Supreme Court.
    Hearing before Supreme Court.
    query: कार्यवाही को अमान्य नहीं करने के लिए अधिनियम में किस धारा के तहत रिक्त पद आदि को कवर किया गया है? passage: धारा 382: रिक्तियां, आदि, कार्यवाही को अमान्य नहीं करने के लिए 382. अग्रिम निर्णय बोर्ड द्वारा कोई कार्यवाही या अग्रिम निर्णय की घोषणा, केवल बोर्ड के गठन में किसी भी रिक्त स्थान या दोष की उपस्थिति के आधार पर प्रश्न या अमान्य होगी। अग्रिम निर्णय के लिए आवेदन। passage: धारा 381: अग्रिम निर्णयों के लिए बोर्ड 381. (1) केंद्र सरकार इस अध्याय के तहत अग्रिम निर्णय देने के लिए एक या अधिक बोर्डों का गठन करेगी, जैसा कि आवश्यक हो, उस तिथि पर या उसके बाद, जिसे केंद्र सरकार अधिसूचना द्वारा नियुक्त कर सकती है। (2) अग्रिम निर्णयों के लिए बोर्ड में दो सदस्य होंगे, प्रत्येक एक अधिकारी मुख्य आयुक्त के पद से कम नहीं होगा, जैसा कि बोर्ड द्वारा नामित किया जा सकता है। रिक्तियां आदि, कार्यवाही को अमान्य नहीं करने के लिए।
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 96
  • learning_rate: 2e-05
  • num_train_epochs: 1
  • lr_scheduler_type: cosine
  • warmup_ratio: 0.1
  • seed: 45
  • dataloader_num_workers: 4

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • prediction_loss_only: True
  • per_device_train_batch_size: 96
  • per_device_eval_batch_size: 8
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 2e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 1
  • max_steps: -1
  • lr_scheduler_type: cosine
  • lr_scheduler_kwargs: None
  • warmup_ratio: 0.1
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 45
  • data_seed: None
  • jit_mode_eval: False
  • bf16: False
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 4
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch_fused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss
0.2381 10 2.4989
0.4762 20 2.3856
0.7143 30 2.3208
0.9524 40 2.343

Training Time

  • Training: 15.6 minutes

Framework Versions

  • Python: 3.12.3
  • Sentence Transformers: 5.7.0
  • Transformers: 4.57.6
  • PyTorch: 2.13.0+cpu
  • Accelerate: 1.14.0
  • Datasets: 5.0.1
  • Tokenizers: 0.22.2

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
33
Safetensors
Model size
33.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for vivekkopthsd/hinglish-st-33m